Google ships Gemini 3.8 Flash TTS and Flash-Lite TTS with prompt-designed voices

Google announced two new text-to-speech models in the Gemini family, saying they turn voice generation "from static presets into a dynamic creative studio". Gemini 3.8 Flash TTS is built for creative direction and character design. Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale: dubbing, audio content creation and expressive voice agents, with fine-grained control over tone, pacing and expressive nuance. Google says they join a Gemini Audio family that already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live and 3.8 Live Extended Thinking.
The headline feature of Flash TTS is voice creation. Instead of picking from 30 original voices, users can design new ones from a natural language prompt, setting role, accent and voice characteristics across more than 100 languages and dialects. Google's examples range from a fire-breathing dragon to a narrator with a distinct regional cadence. Alongside that sit a library of 2,000+ production-ready voices (including regional varieties such as Mexican Spanish, Quebec French and Scots English), voice replication from a 30-second audio sample of your own voice or one you have the rights to use, and a save-and-manage feature so custom voices stay consistent across projects with minimal drift. Voice remixing, where you pick a library voice and adjust timbre, pitch, pace and accent with prompts such as "add subtle Southern US accent", is labelled "coming soon".
Both models give line-by-line control over delivery. Users can write their own stage directions or let Gemini steer delivery from script cues, from a calm customer service agent to a whispered suspense scene. Google also lists long-form generation that holds voice quality, pacing and character timbre across hours of continuous audio with minimal speaker drift; native two-speaker scene staging from a single script, keeping both voices distinct with natural turn-taking; and scripted vocal bursts and backchanneling, using cues like
On quality, Google says Gemini 3.8 Flash TTS took the #1 overall spot on Hume AI's Voice Design Benchmark with a score of 71.4 and leads in accent modeling with 60.8. It says Flash TTS and Flash-Lite TTS take #1 and #2 respectively on Hume AI's Overall Quality Index. Compared with Gemini 3.1 Flash TTS, it describes major improvements on long-form content and dual-speaker screenplay control. In blind human preference evaluations on Voice Arena, Google says both models hold top positions among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. The models support over 100 languages.
Google describes safeguards for voice creation and replication. To replicate a voice, users must provide a verbal consent recording from the voice owner that matches the reference speaker before the voice can be created. Replication is also backed by SynthID watermarking and C2PA credentials, and Google says every audio clip generated by its Gemini Audio models carries a SynthID watermark. It points readers to the model card for more on safety.
Availability: starting today, developers can use both models in the Gemini API and Google AI Studio, where a new audio playground works like a voice design workspace. You can prompt new vocal identities or replicate your own voice, then move them into a dual-speaker screenplay editor to direct delivery line by line. Consumers get Flash TTS in Gemini Notebook and Flash-Lite TTS in Google Vids. Enterprise access through the API in Gemini Enterprise is "coming soon" for both. Developer platforms Agora, LiveKit, Pipecat and Vercel build on the Gemini API. Google says it is partnering with Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang, who are integrating the models for global dubbing, localized media with regional accents and conversational voice agents at scale.
Key facts
- Google introduced Gemini 3.8 Flash TTS (creative direction and character design) and Gemini 3.8 Flash-Lite TTS (high-volume, cost-efficient work), rolling out from the day of the post.
- Flash TTS can create new voices from natural language prompts across more than 100 languages and dialects, and offers a 2,000+ voice library and replication from a 30-second sample.
- Google reports Flash TTS at #1 on Hume AI's Voice Design Benchmark (71.4) and accent modeling (60.8), and #1 and #2 for Flash and Flash-Lite on Hume AI's Overall Quality Index.
- Replication requires a verbal consent recording from the voice owner, and every clip from Gemini Audio models is watermarked with SynthID.
- Developers get both models in the Gemini API and Google AI Studio; Flash TTS reaches consumers in Gemini Notebook, Flash-Lite TTS in Google Vids; Gemini Enterprise and voice remixing are coming soon.
Why it matters
Google frames the release as a move from fixed voice presets to designing voices from a text prompt, at a scale of more than 100 languages and dialects. The models extend the Gemini Audio family, which already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live and 3.8 Live Extended Thinking. Google also claims improvements over Gemini 3.1 Flash TTS in long-form content and dual-speaker screenplay control, though it gives no figures for that comparison.
Who it affects
Google names creators, developers and enterprises. The use cases it lists are gaming, audiobooks, podcasts, interactive media, dubbing and voice agents. Voice talent is also affected, since the models can replicate a voice from a 30-second sample. Google names Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as partners integrating the models, and Agora, LiveKit, Pipecat and Vercel as developer platforms building on the Gemini API.
How to use it
Developers can start today in the Gemini API and Google AI Studio. The AI Studio audio playground lets you prompt new vocal identities or replicate your own voice, then direct delivery in a dual-speaker screenplay editor. Direction can be written as your own stage directions or left to script cues, with tags such as
How solid is it
This is Google's own announcement, and every performance claim comes from Google. The benchmark figures are 71.4 on Hume AI's Voice Design Benchmark and 60.8 in accent modeling, but the scale and competitors' scores are not given. Voice Arena results are described only as top positions in named languages, with no scores or ranks. Independent verification of the benchmark claims is not mentioned.
Risks and caveats
Voice replication is the sensitive part. Google's safeguards are a verbal consent recording from the voice owner that must match the reference speaker, plus SynthID watermarking and C2PA credentials on replicated voices, and it says the SynthID watermark is meant to keep AI speech detectable. Those are Google's descriptions of its own system. Some features are not live yet: voice remixing and Gemini Enterprise access are both labelled coming soon, with no dates. No latency, cost or speed figures are given for Flash versus Flash-Lite, so the difference between them is described only in positioning.
“transforming voice generation from static presets into a dynamic creative studio”
— Google, announcement of Gemini 3.8 text-to-speech