AI Audio · 1 Oct 2026 · 23:53 CEST
Microsoft Launches MAI-Transcribe-2-Streaming and Two MAI-Voice Models

Publisher preview · OZZZER analysis pending editorial review.
Microsoft AI on October 1, 2026 launched MAI-Transcribe-2-Streaming, its first streaming transcription model, alongside two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, with all three available through Microsoft Foundry. Microsoft describes MAI-Transcribe-2-Streaming as delivering low-latency, real-time transcripts in 60 languages with automatic, continuous language detection. The company said the model ranks No. 1 for accuracy for both final and partial transcripts on Artificial Analysis, and that it sits on the Pareto frontier of the benchmark’s accuracy-versus-latency evaluation, meaning higher accuracy does not require a heavy latency tradeoff.
The leaderboard chart in the post, citing the Artificial Analysis streaming leaderboard dated September 28, 2026, shows the model at a 2.5 percent final word-error rate, a 2.8 percent first-partial rate, and 0.13 seconds to final transcription. Rather than waiting for a speaker to finish before returning text, the model produces its first hypotheses, known as partials, in just over 100 milliseconds of receiving audio, then revises them as more context arrives before committing a stable transcript.
Microsoft said this allows voice-enabled applications to act on speech before the speaker finishes: voice agents can…
Excerpt supplied by the publisher.
Source
Unite.AI · 1 Oct 2026 · 23:53 CEST
Open the original at Unite.AI ↗