Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.
153 comments
just remember: despite they would have you believe they are are a united front... they are better thought of as a "loose federation of warring tribes".
I’m reducing into silliness I’m sure, I don’t have any breadth or depth of knowledge here.
Edit: wonder why Apple is ostensibly different. MS seems similar, and don’t know enough about AWS—maybe I’ve seen complaints about them being disjointed, but not as much as Google & Microsoft.
Googly way to put it! Nice. Might steal that :)
My personal account, 3.6 Flash Lite and 3.5 Thinking.
Meanwhile, I can go hog wild and drain my bank account on GCP. I don’t though, because my family has to eat.
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
After some debugging, making a clean dataset with clean recordings, and experimenting with a good fine tune recipe (much props to the new GPT models yesterday being cheaper).
I was able to make a robo-me that sounds absurdly good, family was shocked, all in a matter of a few hours.
So yeah the cat is out of the bag for sure.
It took like 1min - capture something from a youtube or video and put in your own text. It worked also really good for a german test.
Made a voice message for my wife from one of our favorite actors, telling here how nice it would be to make some breakfast :D
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
Probably the latter. Cat's already out of the bag to the extent that you can synthesize with a specific voice in one go and it sounds decent. Even if you need commercial models for better intonation or whatever, you can probably get the commercial models to first generate with a generic voice, then use a local model to transfer that to voice you're cloning. That'll probably get rid of any C2PA watermarks too.
https://www.youtube.com/watch?v=WAeHgE94rVo
No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.
[1]: https://deepmind.google/models/gemma/gemma-4/
[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design
I wasn't able to find a version of these that can create voice samples based on voice designs. Do you mean to use Qwen3 TTS Voice Design to create samples followed by Higgs or Fish Audio to clone the sample voices and narrate the novel?
MOSS-TTS 2.0 will apparently have voice design, as well, on par with ElevenLabs quality.
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion
Is it possible to annotate your text with extra 'stage directions' that influence how the book is read out?
Good idea, not something I've considered yet. Wouldn't take much to add it since there's already a feature for selecting a quotation and assigning it an intonation. Same infrastructure could be reused to select arbitrary text and assign stage directions.
The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.
Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.
So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.
Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS
I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/
I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!
My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...
Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..
I found all these on this subreddit: https://www.reddit.com/r/TextToSpeech/
I've also see comments like "Microsoft Edge's read-aloud feature is amazing for TTS" but I haven't tried it myself.
Personally I use Kokoro with a python front-end on my Macbook, I linked to it in this comment: https://news.ycombinator.com/item?id=49818923 it outputs MP3, so I just copy those to my phone and listen as audiobooks.
Please let me know if you have any questions or feedback.
It sends it to your Spotify playlist so you can listen in the same place as the rest of your podcasts etc
It was especially nice during a bike trip along the Rhine, I listened to a lot of the history of the industrial area and its cities.
Read the full thread on Hacker News →
Related stories
- DEV Community · 2 points · 4 days ago
- Ars Technica · 0 points · 9 days ago
- DEV Community · 1 points · 3 days ago
- DEV Community · 4 points · about 7 hours ago
- Ars Technica · 0 points · about 7 hours ago
- Hacker News · 1 points · 1 day ago