Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS

Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS

Google introduced two new text-to-speech models to the Gemini family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The company frames the release as turning voice generation from static presets into what it calls a dynamic creative studio. Flash TTS is built for deep creative direction and character design, letting creators build entirely new voices from natural language prompts and direct every line of a performance with control over acting cues, pacing, dialect shifts and backchanneling, aimed at gaming, audiobooks, podcasts and interactive media. Flash-Lite TTS is built for high-volume, cost-efficient use such as dubbing, audio content creation and expressive voice agents, with fine-grained control over tone, pacing and expressive nuance. Both join Google's Gemini Audio family alongside 3.5 Live Translate, 3.5 Transcribe, 3.8 Live and 3.8 Live Extended Thinking.

On voice creation, Google says the update scales up from 30 original preset voices to what it describes as an infinite library. Generative voice design with Flash TTS lets users create bespoke voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting. An expansive voice library offers 2,000 or more production-ready voices with broad language coverage, including regional varieties such as Mexican Spanish, Quebec French and Scots English. A voice replication feature can recreate a consistent vocal profile from just a 30-second audio sample of a person's voice, or a voice the user has the rights to use, backed by consent verification, SynthID watermarking and C2PA credentials. A separate voice remixing feature, letting users pick a voice from the library and fine-tune timbre, pitch, pace and accent through prompts, is described as coming soon rather than available now.

Both TTS models give precise control over how each line is delivered: users can write their own stage directions or let Gemini interpret natural script cues, from a calm customer-service agent to a whispered suspense scene. The models are built to maintain voice quality, natural pacing and character timbre across hours of continuous audio with minimal drift, aimed at podcasts and audiobooks. A native two-speaker scene-staging feature can direct multi-turn conversations from a single script while keeping voices distinctly separated with natural turn-taking, and scripted vocal bursts and backchanneling (non-verbal cues like , , , and interjections like 'mhm' or 'yeah') add conversational texture.

On quality, Google says Gemini 3.8 Flash TTS took the number-one overall spot on Hume AI's Voice Design Benchmark with a score of 71.4, and also led in accent modeling with a score of 60.8. Flash TTS and Flash-Lite TTS took the number-one and number-two spots, respectively, on Hume AI's Overall Quality Index, with Google citing major improvements over Gemini 3.1 Flash TTS on long-form content and dual-speaker screenplay control. In blind human preference evaluations on Voice Arena, the two models placed at the top among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi, with support claimed for over 100 languages overall.

On safety, Google says voice replication requires consent verification: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Every audio clip generated by Gemini Audio models is watermarked with SynthID, an imperceptible watermark Google says keeps AI-generated speech detectable to help prevent misinformation.

Both models are rolling out starting today. Gemini 3.8 Flash TTS is available to developers through the Gemini API and Google AI Studio, and to consumers in Gemini Notebook; enterprise access via the Gemini Enterprise API is coming soon. Gemini 3.8 Flash-Lite TTS is available through the same developer channels and to consumers in Google Vids, with enterprise API access also coming soon. Google AI Studio now includes an audio playground where developers can prompt new vocal identities from scratch or replicate their own voice, then bring them into a dual-speaker screenplay editor. Developer platforms including Agora, LiveKit, Pipecat and Vercel are cited as enabling teams to build and deploy speech-generation experiences using the Gemini API, and Google names Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as partners integrating the new TTS models to help with dubbing, localization and conversational voice agents. No pricing was disclosed for either model.

Key facts

  • Google introduced Gemini 3.8 Flash TTS, built for creative voice design and line-by-line performance direction, and Gemini 3.8 Flash-Lite TTS, built for high-volume, cost-efficient dubbing and voice agents.
  • The voice library scales from 30 original preset voices to 2,000+ production-ready voices, with generative voice design spanning more than 100 languages and dialects via natural language prompting.
  • Voice replication can recreate a consistent voice from a 30-second sample, gated by mandatory verbal consent verification, SynthID watermarking and C2PA credentials; a voice remixing feature is coming soon but not yet available.
  • Gemini 3.8 Flash TTS took the number-one spot on Hume AI's Voice Design Benchmark (71.4 overall, 60.8 in accent modeling); Flash TTS and Flash-Lite TTS placed first and second on Hume AI's Overall Quality Index.
  • Both models roll out starting today via the Gemini API and Google AI Studio for developers, plus Gemini Notebook and Google Vids for consumers; enterprise access via Gemini Enterprise is coming soon.

Why it matters

Google is repositioning Gemini's text-to-speech tools from a set of fixed preset voices into a generative voice-design studio: users can prompt entirely new voices into existence, clone a voice from a short sample, and direct line-by-line delivery rather than picking from a limited menu. Google backs the move with benchmark claims, citing the number-one spot on Hume AI's Voice Design Benchmark and top rankings on its Overall Quality Index, positioning the release as a quality leap over the prior Gemini 3.1 Flash TTS generation.

Who it affects

The release targets creators (game studios, audiobook and podcast producers), developers building voice agents and dubbing pipelines, and enterprises doing media localization, as well as everyday users through Gemini Notebook and Google Vids. Google names developer platforms Agora, LiveKit, Pipecat and Vercel as enabling Gemini API voice deployments, and lists Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang as partners integrating the new models for dubbing, localization and conversational agents.

How to use it

Both models are rolling out starting today. Developers get Gemini 3.8 Flash TTS and Flash-Lite TTS through the Gemini API and a new audio playground in Google AI Studio, where they can design voices from prompts, clone a voice from a sample, and edit dual-speaker scripts line by line. Consumers get Flash TTS inside Gemini Notebook and Flash-Lite TTS inside Google Vids. Enterprise access to both models via the Gemini Enterprise API is coming soon, as is a voice-remixing feature for fine-tuning an existing library voice's timbre, pitch, pace and accent through prompts. No pricing was disclosed.

How solid is it

This account is based on Google's own product announcement, including all benchmark figures (Hume AI's Voice Design Benchmark and Overall Quality Index, and the Voice Arena human-preference evaluations); no independent verification of those scores appears in the source, and the post does not name an individual spokesperson or author.

Risks and caveats

Voice replication from a short sample raises obvious impersonation risk, which Google says it addresses by requiring a verbal consent recording from the voice owner before a cloned voice can be created, alongside SynthID watermarking and C2PA credentials on generated audio; the source does not detail how consent verification is technically enforced. Several announced capabilities, including voice remixing and enterprise API access via Gemini Enterprise, are described as coming soon rather than available at launch.