Gemini 3.8 Flash TTS and a lighter sibling, Gemini 3.8 Flash-Lite TTS, move Google's speech synthesis from a short menu of stock voices toward voices a developer describes in plain language. Google DeepMind is rolling both Gemini 3.8 models out now in the Gemini API and Google AI Studio. Business customers wait longer, since Gemini 3.8 access through Gemini Enterprise is listed as coming later. Each Gemini 3.8 model targets a different workload.
Google DeepMind pitches the Flash version at character work such as game scripts, audiobooks and podcasts, where a writer specifies role, accent and vocal traits in a prompt across more than 100 languages and dialects. Flash-Lite aims at large dubbing jobs and voice agents, where volume and price outweigh fine creative control. Everyday users meet them in different places: the Flash model powers speech in Gemini Notebook, while Flash-Lite shows up in Google Vids.
The voice catalog expands sharply. Google DeepMind says developers go from 30 original voices to a library of more than 2,000 ready-made ones, including regional varieties such as Quebec French, Mexican Spanish and Scots English. Voice replication builds a reusable profile from a 30-second recording, and designed voices can be saved so a character sounds the same across a long project. A remix tool for adjusting the timbre, pitch, pace and accent of library voices is promised but not yet available.
Direction happens inside the script. Writers can attach stage notes to individual lines, stage two speakers from a single script, and drop in tags for laughs, sighs, gasps and listener interjections like "mhm" so an exchange sounds less mechanical. Google also claims the Flash model keeps a voice's character steady across hours of continuous audio, the property that decides whether synthetic narration holds up over a full audiobook.
On quality, Google points to Hume AI's Voice Design Benchmark, where it says the Flash TTS model ranks first overall with a score of 71.4 and leads the accent modeling category at 60.8. The company adds that Flash and Flash-Lite hold the top two positions on Hume AI's Overall Quality Index. Those rankings come from Google's own materials, not from an independent test.
Cloning a voice requires a spoken consent recording from the voice's owner that matches the reference speaker, and Google says every clip from its Gemini audio models carries an imperceptible SynthID watermark. Voice replication in AI Studio is not offered in Illinois, Texas, the EEA, the UK, Switzerland or India, so a meaningful share of developers will get voice design without cloning. With Figma, HeyGen and Wondercraft named as early integrators, and platforms such as LiveKit and Pipecat reaching the models through the Gemini API, the new voices are likely to surface inside other companies' products before most people hear them in Google's own apps.













