Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking for voice agents

Image: Google DeepMind
Why it mattersA voice agent that keeps talking while it runs tools in the background changes what the caller experiences during a slow API call, and a team building phone or in-app voice can now test that shape instead of designing around silences.
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 for building voice agents. Both models handle real-time voice dialogue, run tools in the background while the conversation continues, switch between 97 supported languages mid-conversation, and process live visual input from a camera or screen. Google says the two models are available today through the Gemini API, Google Workspace and the Gemini app.
Two models, one built for scale and one for reasoning
Gemini 3.8 Live is the smaller and cheaper of the two. Google positions it for high-volume voice work and says it took second place on the Speech Agent Arena, the public leaderboard where users compare voice models blind.
Gemini 3.8 Live Extended Thinking is the reasoning tier. It reasons and speaks at the same time, using early verbal cues like "let me check that" to acknowledge a request while the model works, and narrates progress through multi-step background tasks as they run. Google reports it scores 82.6 on the Artificial Analysis Speech to Speech Quality Index, which the company says is the top overall score on that benchmark.
The benchmark numbers Google reports
Google published four numbers for Extended Thinking against three third-party benchmarks. On τ-Voice, a task completion benchmark for voice agents, it reports 68.6 percent. On Sierra's τ-Voice-banking, a harder domain-specific version, it reports 35.1 percent. On Big Bench Audio, a reasoning benchmark, it reports 97.7 percent. All four are Google's own reported figures on external benchmarks, not independent reruns.
Google also shows a chart from ServiceNow's EVA-Bench, which the company says its models "push the Pareto Frontier" on. Google notes the EVA-Bench run was carried out on the Live API through Gemini Enterprise Agent Platform, not on the raw model.
What is actually different about running tools in the background
The specific behaviour Google highlights is that a call to a tool or an external API no longer stops the conversation. The model acknowledges the request out loud, keeps talking, and returns to the tool's result when it arrives. Google demos this with a multi-step booking and with the model turning a hand-drawn sketch and a spoken description into a React component while narrating what it is doing.
For visual input, Google says both models process camera or screen frames in near real time and use that context in their reply. The employee-onboarding demo has the model looking at what a new hire is seeing on their laptop while answering questions about it.
The route into it
The models are in the Gemini API today, in Google Workspace inside Docs Live, Gmail Live and Keep Live, and in Search Live, which Google separately confirmed is now powered by Gemini 3.8 Live. Google did not publish per-token pricing in the launch post, only that the models are "highly cost-effective" and "cost efficient" relative to other frontier speech models. A team costing a voice product should read the API pricing page rather than trust that phrase.
A voice agent that fills silence with a real acknowledgement and progress narration is a different call experience than one that goes quiet for eight seconds while a database query runs. Whether that shows up in the completion rate on your own tasks is the number worth measuring locally before adopting it.
Source
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, Tom Ouyang and Malini Jaganathan, Google DeepMind, 15 September 2026.
Source: Google DeepMind
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually shipping with them. Short, and only when there is something worth reading.


