Models & agents

Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking for voice agents

September 15, 2026 at 5:30 PM PT

Google DeepMind blog header for Gemini 3.8 Live and 3.8 Live Extended Thinking

Image: Google DeepMind

Why it mattersA voice agent that keeps talking while it runs tools in the background changes what the caller experiences during a slow API call, and a team building phone or in-app voice can now test that shape instead of designing around silences.

Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 for building voice agents. Both models handle real-time voice dialogue, run tools in the background while the conversation continues, switch between 97 supported languages mid-conversation, and process live visual input from a camera or screen. Google says the two models are available today through the Gemini API, Google Workspace and the Gemini app.

Two models, one built for scale and one for reasoning

Gemini 3.8 Live is the smaller and cheaper of the two. Google positions it for high-volume voice work and says it took second place on the Speech Agent Arena, the public leaderboard where users compare voice models blind.

Gemini 3.8 Live Extended Thinking is the reasoning tier. It reasons and speaks at the same time, using early verbal cues like "let me check that" to acknowledge a request while the model works, and narrates progress through multi-step background tasks as they run. Google reports it scores 82.6 on the Artificial Analysis Speech to Speech Quality Index, which the company says is the top overall score on that benchmark.

The benchmark numbers Google reports

Google published four numbers for Extended Thinking against three third-party benchmarks. On τ-Voice, a task completion benchmark for voice agents, it reports 68.6 percent. On Sierra's τ-Voice-banking, a harder domain-specific version, it reports 35.1 percent. On Big Bench Audio, a reasoning benchmark, it reports 97.7 percent. All four are Google's own reported figures on external benchmarks, not independent reruns.

Google also shows a chart from ServiceNow's EVA-Bench, which the company says its models "push the Pareto Frontier" on. Google notes the EVA-Bench run was carried out on the Live API through Gemini Enterprise Agent Platform, not on the raw model.

What is actually different about running tools in the background

The specific behaviour Google highlights is that a call to a tool or an external API no longer stops the conversation. The model acknowledges the request out loud, keeps talking, and returns to the tool's result when it arrives. Google demos this with a multi-step booking and with the model turning a hand-drawn sketch and a spoken description into a React component while narrating what it is doing.

For visual input, Google says both models process camera or screen frames in near real time and use that context in their reply. The employee-onboarding demo has the model looking at what a new hire is seeing on their laptop while answering questions about it.

The route into it

The models are in the Gemini API today, in Google Workspace inside Docs Live, Gmail Live and Keep Live, and in Search Live, which Google separately confirmed is now powered by Gemini 3.8 Live. Google did not publish per-token pricing in the launch post, only that the models are "highly cost-effective" and "cost efficient" relative to other frontier speech models. A team costing a voice product should read the API pricing page rather than trust that phrase.

A voice agent that fills silence with a real acknowledgement and progress narration is a different call experience than one that goes quiet for eight seconds while a database query runs. Whether that shows up in the completion rate on your own tasks is the number worth measuring locally before adopting it.

Source

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, Tom Ouyang and Malini Jaganathan, Google DeepMind, 15 September 2026.

Source: Google DeepMind

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

More from AI News

Semrush ran a manufacturing AI-search study on 458 keywords and found only two brands appear on both the most-mentioned and the most-cited top-15 lists

Semrush tracked 458 manufacturing keywords from January to July 2026 across ChatGPT, Gemini, Google AI Mode, and AI Overviews, and reports that mentioned brands and cited sources are almost different lists, with only Vevor and Grainger appearing on both top-15.

Source: PressGo-to-market

Amazon Science raises LLM-as-a-judge accuracy by 9 to 14 points by modelling how the judges copy each other

Amazon Science reports that panels of LLM judges often agree because they share training lineage or prompt templates, and that modelling those correlations with an Ising model raises aggregation accuracy by 9 to 14 points across three tasks.

Source: Hacker NewsModels & agents

IBM Research measures a 24-point consistency gap on AppWorld and halves it with automatic guidelines

IBM Research reports that a ReAct agent posting 77.4 percent on the AppWorld benchmark succeeds on all five repeated runs for only 53 percent of tasks, and that a diagnostic plus targeted guidelines cut the 24-point gap in half without hurting average accuracy.

Source: Vendor blogModels & agents