signalSimon Willison's Blog2026-09-24
Gemini 3.8 TTS Playground
Google released two new text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, with over 2,000 voices and custom voice creation from a 30-second audio sample. A playground interface was built with GPT-6 Astra and a bring-your-own-key setup, leveraging the API's open CORS policy. Generating 1 minute 18 seconds of audio took about 20 seconds and cost 2.74 cents on Flash TTS.
- for who
- Developers and creators who need multi-voice, multi-character text-to-speech generation
- why now
- Google's new Gemini TTS models, released today, offer 2,000 voices and 30-second custom voice cloning.
- what changes
- They can now produce full conversations with distinct voices and style instructions without complex audio editing
- to do
- Try the playground with your own Gemini API key and create a custom voice from a 30-second sample
key points
- Two new Gemini TTS models released, flash and flash-lite
- Over 2,000 voices plus custom voice from 30-second audio sample
- 1m18s audio generated in 20 seconds for 2.74 cents
#gemini#tts#multi-character dialogue
score
score 9 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source