Models & tools · 3 Oct 2026 · 07:09 CEST
Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis
Publisher preview · OZZZER analysis pending editorial review.
Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy. The model targets voice agents, live captions and dictation, where latency decides the experience. What Microsoft Shipped MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-Transcribe-2, released in September.
It transcribes 60 languages with automatic, continuous language detection. Audio streams in continuously, and text streams back while the speaker is still talking. The model emits its first hypotheses, called partials, just over 100ms after receiving audio. It revises those partials as context arrives, then commits a stable final transcript. An agent can therefore start reasoning or calling tools mid-sentence.
Microsoft team states its internal tests show words appearing 2x faster than its closest competitor. What Artificial Analysis Measured The AA-WER Streaming index uses about 8 hours of audio. The mix is AA-AgentTalk (50%), VoxPopuli (25%) and Earnings22 (25%). Latency is timed from the end of speech, as detected by SileroVAD. Final…
Excerpt supplied by the publisher.
Source
MarkTechPost · 3 Oct 2026 · 07:09 CEST
Open the original at MarkTechPost ↗