AI NewsModels & agentsAnnouncement
Microsoft released three audio models with transcription at 54 cents an hour and voice at 15 dollars per million characters
Microsoft published three audio models on 1 October 2026, pricing streaming transcription at $0.54 per hour of audio and MAI-Voice-2.1-Flash at $15 per million characters, and claiming first partial transcripts in just over 100 milliseconds.

Image: Microsoft AI
Why it mattersA product adding voice turns to a model is now looking at hourly transcription that costs less than a dollar and sub-200ms voice replies, which is the price and latency band where read-aloud answers and spoken agents start to make sense on every user interaction.
The reason most products do not read their replies back out loud is price. Microsoft published three audio models on 1 October 2026 that cut that price and the latency behind it.
The company's blog post titled "Our first streaming transcription model" names the three. MAI-Transcribe-2-Streaming is a speech-to-text model priced at $0.54 per hour of audio, an introductory rate that runs through the end of 2026. MAI-Voice-2.1 is the full text-to-speech model, priced at $22 per million characters. MAI-Voice-2.1-Flash is a faster, lower cost version at $15 per million characters.
The latency and accuracy claims
Microsoft says MAI-Transcribe-2-Streaming returns a first partial transcript in "just over 100ms" and that words appear in the transcript "2x faster than with our closest competitor". The post does not name that competitor. The company ranks the model top on Artificial Analysis for both final and partial transcript accuracy, and says it supports 60 languages with automatic continuous language detection.
MAI-Voice-2.1-Flash is claimed to produce 45 seconds of audio with end to end latency of 150 milliseconds. Microsoft says the Flash variant is "55% faster model inference" and "~60% cheaper than comparable models", again without naming the models in the comparison. Both voice models cover 23 languages and 26 locales and support voice cloning from a few seconds of reference audio, with one voice holding its accent across every language it speaks.
Where you can call them
The announcement lists five routes. The models are available on Microsoft Foundry, the MAI Playground, OpenRouter, Vercel, and Azure Voice Live, with LiveKit integration listed as coming soon. OpenRouter and Vercel put the pricing inside the same API gateway many AI products already route through, which lowers the setup cost of trying them against the ElevenLabs or OpenAI call a product makes today.
The two numbers that held voice replies back
A product adding a voice reply on every user interaction has been held off by two numbers: a per character price in the tens of dollars for a million characters, and a round trip that sits above a second once the model, the network and the audio pipeline are added together. $15 per million characters at 150 milliseconds end to end is below the band where read aloud answers were worth adding, which is why Microsoft's pricing matters even to a team that is currently happy with its vendor.
The claims are Microsoft's own. Artificial Analysis is a third party ranking, so the top accuracy position can be checked against that site. The "2x faster" and "60% cheaper" numbers are measured by Microsoft against unnamed competitors, and a team evaluating these would want to run its own transcript against the current provider before switching, especially on a language the current vendor is strong in.
Source
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.
