Microsoft's MAI-Transcribe-2-Streaming Tops 38 Models With 2.5% Error Rate

Microsoft AI's new streaming speech-to-text model tops the Artificial Analysis leaderboard with 2.5% WER delivered just 0.13 seconds after speech ends.

·
·
Microsoft's MAI-Transcribe-2-Streaming Tops 38 Models With 2.5% Error Rate
  • Microsoft AI released MAI-Transcribe-2-Streaming, taking #1 on Artificial Analysis streaming WER.
  • Hits 2.5% WER on final transcript at 0.13s after end of speech, beating Grok Voice Transcribe 2.0 at 2.7%.
  • Also leads First Partial Transcript at 2.5% WER and 0.12s latency, ahead of ElevenLabs and Muse.
  • Priced at $0.54 per audio hour for streaming, matching Gemini but above ElevenLabs and Deepgram Flux.
  • Non-streaming MAI-Transcribe-2 stays at a promotional $0.10 per hour through end of 2026.
  • Covers 60 languages with diarization, word-level timestamps, keyword biasing, available on Azure Speech preview.

Microsoft’s streaming transcriber leads an independent benchmark

Microsoft AI has released MAI-Transcribe-2-Streaming, a real-time counterpart to its batch transcription model. At publication, Artificial Analysis ranks it first among 38 models for final transcripts, with a 2.5% word error rate and 0.13 seconds of post-speech latency. Low error rates reduce corrections, while short delays help captions and voice agents respond sooner.

The release extends a line that includes MAI-Transcribe-1.5 and the batch MAI model. Streaming recognition emits partial text while audio arrives. The batch edition processes completed recordings. Reports on the project history and team structure attribute the models to a ten-person core group supported by larger data and vendor teams.

Lowest WER, near-fastest response

Artificial Analysis evaluates recognition quality through word error rate, or WER, alongside the time required to produce partial and final transcripts. Lower WER is better: a 2.5% score represents roughly 2.5 substitutions, deletions, or insertions per 100 reference words.

Leading final-transcript results
Model Final WER Latency after speech ends
MAI-Transcribe-2-Streaming 2.5% 0.13 s
Grok Voice Transcribe 2.0 2.7% 0.49 s
Muse Voice Transcribe 3.1% 0.16 s
Cartesia Ink Preview 3.1% 0.11 s
ElevenLabs Scribe v2 Realtime 3.6% 0.14 s

In the first-partial view, MAI pairs its 2.5% WER with 0.12 seconds to the first partial transcript. Grok Voice Transcribe 2.0 records 3.4% at 0.49 seconds. Cartesia Ink-2 responds faster at 0.07 seconds, with a higher 4.0% WER. No listed model improves on both of MAI’s figures simultaneously.

Production results can shift with language, accent, background noise, overlapping speakers, domain vocabulary, endpointing rules, and network distance. Partial-transcript stability also matters because repeated revisions can cause interface flicker or trigger an agent too early. The leaderboard provides a useful screening result, while representative audio remains the decisive test.

The live model carries a premium

MAI-Transcribe-2-Streaming costs $0.54 per audio hour, equivalent to $9.00 per 1,000 minutes. That matches Gemini 3.5 Transcribe Live and exceeds several nearby competitors.

Listed streaming transcription prices
Model Price per 1,000 minutes
MAI-Transcribe-2-Streaming $9.00
Gemini 3.5 Transcribe Live $9.00
ElevenLabs Scribe v2 Realtime $6.50
Deepgram Flux $6.50
Cartesia Ink-2 $4.00
Muse Voice Transcribe $3.00

Microsoft lists the batch MAI-Transcribe-2 model at a promotional $0.10 per audio hour through the end of 2026. MAI-Transcribe-1.5 costs $0.36 per hour. At the listed rates, 1,000 audio hours would cost $100 with the promotional batch model and $540 with the streaming edition, before surrounding infrastructure costs.

Preview APIs shape deployment

Access paths for the MAI Transcribe models
Workload Model Access Status
Prerecorded audio MAI-Transcribe-2 Microsoft Foundry, MAI Playground, OpenRouter, and Azure Speech Azure Speech public preview in East US, West US, Southeast Asia, and North Europe
Live audio MAI-Transcribe-2-Streaming Voice Live API Public preview without an SLA

Prerecorded jobs use the Fast Transcription API, while live audio uses the separate Voice Live API. Microsoft advises against production deployment of the Voice Live preview because it carries no service-level agreement. Customer-facing systems would need a fallback provider or acceptance of preview-level availability.

Microsoft lists the following capabilities for the MAI Transcribe family, although availability can vary by endpoint:

  • One multilingual model covering 60 languages
  • Speaker diarization
  • Word-level timestamps
  • Keyword biasing for names and domain terms
  • Verbatim or cleaned-up output

Where 150 milliseconds helps

The reported sub-150-millisecond response times suit interactions where visible or downstream text must arrive quickly:

  • Live captions: Meetings, lectures, broadcasts, and accessibility tools can display text with less delay.
  • Voice agents: Final transcripts can reach the language model sooner after a speaker finishes.
  • Contact centers: Coaching, compliance checks, and agent assistance can run during calls.
  • Dictation: Partial text can appear while the user continues speaking.

The batch model suits meeting archives, recorded calls, podcasts, media libraries, and other workloads without an interactive latency requirement. Its promotional rate makes large backfills inexpensive, provided the application can wait for prerecorded processing.

Test the full audio path

  1. Audio mix: Evaluate representative languages, accents, microphones, noise levels, overlapping speech, and specialized vocabulary.
  2. End-to-end latency: Measure capture, network transit, endpoint detection, transcription, and downstream processing rather than relying on model latency alone.
  3. Partial stability: Track how often interim text changes and whether those revisions affect captions or agent triggers.
  4. API behavior: Confirm concurrency limits, session duration, regional availability, retry handling, and rate limits.
  5. Feature support: Verify diarization, timestamps, keyword biasing, and output normalization on the chosen endpoint.
  6. Operations and cost: Check current pricing, data retention, residency requirements, logging, fallback behavior, and the absence of a preview SLA.

Microsoft fills the live-audio gap

MAI-Transcribe-2-Streaming gives Microsoft a high-accuracy, low-latency option for live audio alongside its cheaper batch model. Its $9-per-1,000-minute price and preview-only access set the main constraints. Production adoption will depend on workload-specific accuracy, full-path latency, regional support, partial-transcript behavior, and Microsoft adding an SLA.

Trending
  • No trending articles

Comments

avatar

Next Reads