Google launches Gemini 3.8 Live models to power voice agents

Google launches Gemini 3.8 Live models to power voice agents

Google introduced two new Gemini models built for voice agents: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. Both are built around near real-time reasoning, which Google says helps voice agents work more reliably and makes talking with AI feel more natural. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, with greater intelligence and multi-step reasoning. Google positions both as building blocks for developers and enterprises building reliable, production-ready voice agents, and says they also make speaking with Gemini more fluid across the Gemini app, Google Workspace and Search.

Google backs its performance claims for Extended Thinking with several benchmark results: it cites the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index, with a score of 82.6, and says the model leads in agentic task completion with 68.6% on the τ-Voice benchmark and 35.1% on Sierra's τ-Voice-banking benchmark. On reasoning, Google reports a 97.7% score on Big Bench Audio. None of these four figures is given for the base Gemini 3.8 Live model. For that model, Google's evidence is qualitative: it says Gemini 3.8 Live took second place in the Speech Agent Arena, describing that as showing 'a high preference among users,' without stating a score or percentage. Separately, on ServiceNow's EVA-Bench, a benchmark for evaluating voice agents, Google says its models push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality; a note beneath that claim specifies the result was measured on the Live API running on the Gemini Enterprise Agent Platform. Google gives no dollar figure for either model: it calls Gemini 3.8 Live 'highly cost-effective' and says Extended Thinking maintains 'a highly competitive price point compared to other frontier models.'

On capabilities, Gemini 3.8 Live processes visual input in near real-time to add context to conversations, automatically detects and switches between 97 supported languages mid-conversation, and can execute tool calls and API requests in the background while continuing to talk, so it can acknowledge a request and keep chatting while the task finishes. Gemini 3.8 Live Extended Thinking reasons and speaks at the same time: it uses early verbal cues such as 'Let me check that…' to acknowledge a prompt naturally, then narrates its progress live as it works through multi-step background tasks.

Google says developer platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents already use the Gemini Live API to let developers build voice-driven interfaces, handling the real-time media streaming infrastructure so developers can focus on the user experience. It also names Salesforce, Genspark and Lumeris as partner companies it says are excited about the new models' latency, fluidity and tool-calling.

All audio the models generate carries a SynthID watermark woven directly into the output, which Google says keeps AI-generated audio detectable and helps prevent misinformation. Both models began rolling out the day the announcement was published, though the announcement itself gives no exact calendar date, referring to the rollout only as 'today' and 'starting today.' Gemini 3.8 Live is available to developers through the Gemini API and Google AI Studio, to enterprises in private preview in Gemini Enterprise with Gemini Enterprise for Customer Experience coming soon, and to the public through Search Live. Gemini 3.8 Live Extended Thinking is available the same way to developers, in private preview in Gemini Enterprise with Gemini Enterprise for Customer Experience and Google Workspace business customers coming soon, and to the public through Gemini Live; Google AI Pro and Ultra subscribers also get it in Docs within Workspace, and all Google AI subscribers get it in Gmail and Keep.

Key facts

  • Google launched two new voice-focused Gemini models: Gemini 3.8 Live, built for scale and cost efficiency, and Gemini 3.8 Live Extended Thinking, built for high-complexity, multi-step reasoning tasks.
  • Gemini 3.8 Live Extended Thinking takes the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6, and scores 68.6% on the τ-Voice agentic benchmark, 35.1% on Sierra's τ-Voice-banking benchmark, and 97.7% on Big Bench Audio; none of these four figures is reported for the base Gemini 3.8 Live model.
  • Gemini 3.8 Live took second place in the Speech Agent Arena, detects and switches between 97 languages mid-conversation, and runs tool and API calls in the background while continuing to talk.
  • On ServiceNow's EVA-Bench, Google says its models push the Pareto Frontier for complex workflows, a result the company explicitly scopes to the Live API running on the Gemini Enterprise Agent Platform.
  • Both models start rolling out today: developers get them through the Gemini API and Google AI Studio, and Extended Thinking also reaches consumers through Gemini Live plus Docs, Gmail and Keep for Google AI subscribers.

Why it matters

Voice is becoming a primary way people interact with AI agents, and Google is now offering two purpose-built models instead of one: a cost-efficient Gemini 3.8 Live for scale, and a more capable Gemini 3.8 Live Extended Thinking for complex, multi-step work. Both are built around near real-time reasoning, aimed at making voice agents reliable enough for production use rather than demos. Google backs Extended Thinking with a #1 ranking on Artificial Analysis' Speech to Speech Quality Index and strong scores on agentic and reasoning benchmarks, positioning it against other frontier voice models on both quality and price.

Who it affects

Developers and enterprises building voice agents through the Gemini API, Google AI Studio and Gemini Enterprise; the developer-platform ecosystem building on the Gemini Live API, which Google names as including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents; partner companies Google names as Salesforce, Genspark and Lumeris; and end users of the Gemini app, Search Live, and Google Workspace apps such as Docs, Gmail and Keep.

How to use it

Gemini 3.8 Live is available to developers through the Gemini API and Google AI Studio, to enterprises in private preview in Gemini Enterprise with Gemini Enterprise for Customer Experience coming soon, and to everyone through Search Live. Gemini 3.8 Live Extended Thinking is available the same way to developers, in private preview in Gemini Enterprise with Gemini Enterprise for Customer Experience and Google Workspace business customers coming soon, and to everyone through Gemini Live; Google AI Pro and Ultra subscribers also get it in Docs within Workspace, and all Google AI subscribers get it in Gmail and Keep. Google gives no dollar figure for either model: it calls Gemini 3.8 Live 'highly cost-effective' and says Extended Thinking maintains 'a highly competitive price point compared to other frontier models.'

How solid is it

Every performance figure in the announcement is Google's own. The #1 Speech to Speech Quality Index spot, the 68.6% and 35.1% agentic-task scores, and the 97.7% Big Bench Audio score all describe Gemini 3.8 Live Extended Thinking specifically, not the base Gemini 3.8 Live, which Google describes only qualitatively, as taking second place in the Speech Agent Arena with no score attached. The EVA-Bench claim that its models 'push the Pareto Frontier for complex workflows' carries its own scope note: Google states the result was measured specifically on the Live API running on the Gemini Enterprise Agent Platform, not as a general claim across configurations.

Risks and caveats

No dollar amount or numeric price figure is given for either model, only qualitative language like 'cost-effective' and 'highly competitive price point.' No explicit calendar date appears in the announcement; timing is expressed only as 'today,' 'starting today' and 'coming soon.' No individual person, executive, researcher or spokesperson is named or quoted; the announcement speaks entirely in Google's institutional voice. And Gemini 3.8 Live's 'high preference among users' rests on a single data point, its second-place Speech Agent Arena ranking, with no score or percentage attached to size the gap.