AI NewsModels & agentsAnnouncement

Google adds Gemini 3.8 Flash text-to-speech to the API with 2,000 voices across 100 languages

Google put Gemini 3.8 Flash TTS and Flash-Lite TTS on the Gemini API on 23 September, with over 2,000 voices, per-line style direction, accent controls across 100 languages, and vocal bursts such as laughs and sighs.

AI News

Editorial3 min read

LinkedInX
Google's Gemini audio metacard, from the Gemini 3.8 text-to-speech blog post

Image: Google

Why it mattersA team building a voice agent used to swap in a separate speech vendor with its own API, billing and voice catalogue, and Google is now bundling the whole thing into the same model call.

Google put Gemini 3.8 Flash text-to-speech and Gemini 3.8 Flash-Lite TTS on the Gemini API and in Google AI Studio on 23 September, with over 2,000 production voices across more than 100 languages, per-line style direction, accent controls and vocal bursts such as laughs and sighs. Google says the Flash model tops Hume AI's Voice Design Benchmark and its overall speech quality ranking.

What changed for developers

The previous generation shipped with 30 voices. This one carries what Google calls "2,000+ production-ready voices with broad language coverage," available through the same Gemini API a developer already calls for text. A developer can direct the model line by line, so a script can carry stage directions such as making one line sadder or keeping the next one upbeat. Accent and style can be set across "100+ languages, dialects and regional accents," and the model can emit non-verbal cues on request: laughs, sighs, and what Google calls "backchanneling" for filler noises in a conversation.

Long-form audio

Google says Gemini 3.8 Flash TTS holds voice quality and pacing "across hours of continuous audio," aimed at podcasts, audiobooks and long-running voice agents. This is where a text-to-speech model most often falls apart in production, because a drift in prosody or a slow slide off the target voice is not visible on a short sample. Google does not publish a specific hours-of-continuous-audio number here, and does not say what the previous generation held.

The benchmark, and who runs it

Google reports 71.4 on Hume AI's Voice Design Benchmark, the top score at time of writing, and cites the top two positions on Hume AI's overall speech quality index for its Flash and Flash-Lite variants. Hume AI is a voice-AI company that publishes both the benchmark and, on the same site, its own competing model. That does not make the number wrong, but it is worth reading as one company's chosen scoreboard rather than an independent audit. Google also cites the Voice Arena leaderboard for "top positions amongst competitors in key global languages," without stating the rank.

What developers can plug it into

Google names Agora, LiveKit, Pipecat and Vercel among the platforms that will integrate the new TTS endpoints, which are the same partners often used to plumb voice agents into telephony and browser-side sessions. Both models are available through the Gemini API and Google AI Studio at launch, with Gemini Enterprise access described as coming soon. The documentation is at aistudio.google.com/docs/speech-generation.

The practical change for a team building voice features is that speech generation moves inside the same model provider as the reasoning, so a request no longer has to cross a second vendor's API, quota and billing. The trade is Google's usual one: the voice catalogue, the pricing and the licence come from the same company that runs your text model, so the switching cost the day the pricing changes is one line in the SDK instead of a rebuild. Whether the audio quality matches a specialised vendor on your material is a listening test each team should run on its own scripts, not a ranking from either side.

Source

SourceGoogle

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX