Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.

331 points•swolpers•8 days ago•153 comments•

153 comments

rcr-anti7 days ago
Pet peeve on Google's AI rollouts: there's no alignment across the three platforms they have, consumer, prosumer, cloud. Scroll to the end of every release, including this one, and you'll see different availabilities. The fun part is the models don't even have the same capabilities across platforms! Omni Flash, last I tried and read the docs, is video and text out on consumer and prosumer but video out only on GCP. So if your org disables consumer and prosumer, like mine, it's a coin flip whether you can use the fancy new models or what they can do.
buredoranna7 days ago
> Pet peeve on Google's AI rollouts: there's no alignment across the three platforms they have...

just remember: despite they would have you believe they are are a united front... they are better thought of as a "loose federation of warring tribes".

Barbing7 days ago
I hear much of what happens there is from PMs who want promotions, which makes me wonder if a goal (like interorg alignment) could be met by freezing their promotions until they Google Meet together & hash it out.

I’m reducing into silliness I’m sure, I don’t have any breadth or depth of knowledge here.

Edit: wonder why Apple is ostensibly different. MS seems similar, and don’t know enough about AWS—maybe I’ve seen complaints about them being disjointed, but not as much as Google & Microsoft.

wamatt7 days ago
>they are better thought of as a "loose federation of warring tribes".

Googly way to put it! Nice. Might steal that :)

busssard7 days ago
most large corporations can be seen like this. Thats why stakeholder management is a full time job in those places.
dekhn6 days ago
they even have monkey knife fights (AKA, "resource distribution and trading")
herval7 days ago
united by a Performance Cycle
xattt7 days ago
An education account I have lists 3.1 Pro, and 3.6 Thinking and Flash as the available models in the app.

My personal account, 3.6 Flash Lite and 3.5 Thinking.

Meanwhile, I can go hog wild and drain my bank account on GCP. I don’t though, because my family has to eat.

Melatonic7 days ago
Seriously. It's very annoying. Add to that the corporate Google Workspace also getting models later
deepfriedbits7 days ago
It is and frankly, it's always been a Google weakness.
dansquizsoft7 days ago
I mean, at this point, they are so far behind the frontier, why even bother trying to work with them at all...
simonw7 days ago
> Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.

I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.

kingstnap7 days ago
Yesterday night I was doing a project with QwenTTS 1.7B.

After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).

I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.

So yeah the cat is out of the bag for sure.

MacNCheese237 days ago
Yeah I was doing that at the beginning of this year with voice samples locally from hollywood-stars with Qwen3-TTS.

It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.

Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D

yieldcrv7 days ago
Its so crazy to me how prevalent bad AI voices are, when local models can do such good AI voices
rpastuszak7 days ago
Any chance you could share a bit more detail? I’d love to try this myself but could use some proven structure / approach.
Multicomp7 days ago
They probably do something similar to GPT-Live where they expect a given voice profile to send them a sample saying 'This is the owner of this voice and I consent for synthetic samples to be made of it'

and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

gruez7 days ago
>and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?

Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.

throwa3562627 days ago
This has been possible for quite a long time (probably 1-2 years). There are multiple open models that can do this quite well. Recent example from my YT feed:

https://m.youtube.com/watch?v=WENMgQE9tws

weird-eye-issue7 days ago
> I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
schainks7 days ago
What's the over/under that Android will roll out a spam filter feature that flags AI voice calls that sound like loved ones?
thangalin7 days ago
Here's a video of my Emotive Audiobook Creator, KeenLore, a locally hosted web app:

https://www.youtube.com/watch?v=WAeHgE94rVo

No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.

Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.

[1]: https://deepmind.google/models/gemma/gemma-4/

[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design

RGS18117 days ago
I've been working on a similar project all year and as a tip, you should try Fish Audio or Higgs as a replacement for Qwen3. Both yield much better prosody and are much easier to listen to for long runs.
thangalin7 days ago
> Fish Audio or Higgs

I wasn't able to find a version of these that can create voice samples based on voice designs. Do you mean to use Qwen3 TTS Voice Design to create samples followed by Higgs or Fish Audio to clone the sample voices and narrate the novel?

MOSS-TTS 2.0 will apparently have voice design, as well, on par with ElevenLabs quality.

loremm7 days ago
It's cool technology and I read a lot of audiobooks, even hundreds of hours of TTS. I feel like my brain can fill in the character voices from the text - on the page it's not like they're different fonts.

I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion

eru7 days ago
Awesome! I had been meaning to build something like this for a while now, but never got around to it.

Is it possible to annotate your text with extra 'stage directions' that influence how the book is read out?

thangalin7 days ago
> annotate your text with extra 'stage directions'

Good idea, not something I've considered yet. Wouldn't take much to add it since there's already a feature for selecting a quotation and assigning it an intonation. Same infrastructure could be reused to select arbitrary text and assign stage directions.

Multicomp7 days ago
<grumble grumble people putting in links they expect you to follow to arbitrary goatse youtube videos for all I know>

The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.

Jordan-1177 days ago
Their first sentence literally tells you it's a video of their app? It's not a mystery-meat link.
Multicomp7 days ago
I direct my own extended daydream Star Trek fanfic (okay, I'm on season 2 episode 17) and recently I looked to see if I could have each scene file be read aloud a la an audiobook or radio drama.

Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.

So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.

exhilaration7 days ago
You might find this interesting, it seems to be exactly what you need to make an audio drama: https://github.com/Finrandojin/alexandria-audiobook

Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS

I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/

I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!

My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...

throwa3562627 days ago
Then I think what this guy is doing with local Star Trek voice cloning is right up your alley:

https://m.youtube.com/watch?v=jDudeaWppSE

ghostbrainalpha7 days ago
Original series or Next Generation?
k12sosse7 days ago
Lower decks obviously
hungryhobbit7 days ago
Please don't say Nu Trek.
seemaze7 days ago
My primary use case for TTS is converting written content (blogs, articles, etc.) in to clips I can listen to on the go.

Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..

exhilaration7 days ago
Browser extension? Try these: http://www.paper2audio.com/ https://freevoicereader.com/ https://chromewebstore.google.com/detail/sza-text-to-speech/... https://chromewebstore.google.com/detail/flow-tts/hknidankih...

I found all these on this subreddit: https://www.reddit.com/r/TextToSpeech/

I've also see comments like "Microsoft Edge's read-aloud feature is amazing for TTS" but I haven't tried it myself.

Personally I use Kokoro with a python front-end on my Macbook, I linked to it in this comment: https://news.ycombinator.com/item?id=49818923 it outputs MP3, so I just copy those to my phone and listen as audiobooks.

goldenjm7 days ago
I'm the Paper2Audio founder. I hope you enjoy using us for text to speech. Our browser extension adds your open tabs to your listening queue.

Please let me know if you have any questions or feedback.

kyrra7 days ago
If you use Chrome on Android, there is a "Listen to this page" option in the menu.

https://support.google.com/chrome/answer/14768725?hl=en

e12e7 days ago
Doesn't actually read the page, though. Or not what you see in the browser. For example right now, when I navigate to "reply" it doesn't read the comment I'm replying to - it's reading the login banner you get visiting a reply link without being logged in.
naimkabir2 days ago
Yep, https://webpod.ai

It sends it to your Spotify playlist so you can listen in the same place as the rest of your podcasts etc

unglaublich7 days ago
I vibed a wikipodcast app that has a wikipedia dump, gemma llm for search and summarization and a tts model for audio gen so I could just ask about topics while on the go, and the app would just go off and tell me stuff out of the Wikipedia database.

It was especially nice during a bike trip along the Rhine, I listened to a lot of the history of the industrial area and its cities.

le-mark7 days ago
That’s really great; an ad hoc tour guide!
janalsncm7 days ago
For on the go, I’ve been using ElevenReader. Technically not free but their free tier has been plenty for me.

Read the full thread on Hacker News →

Related stories