Microsoft's in-house AI lab just dropped MAI-Transcribe-1.5, a speech-to-text model that makes a compelling case for a claim the industry rarely gets to make: you don't have to choose between fast and accurate. The model lands at #3 on the Artificial Analysis Word Error Rate (AA-WER) leaderboard with a 2.4% error rate, while simultaneously being the fastest model in that top-10 accuracy tier by a wide margin.

Lapping the competition on speed

MAI-Transcribe-1.5 stands out as the fastest STT model in the top 10 for accuracy, processing audio at ~276x real-time -- more than double the speed of the second fastest model in that accuracy tier. To put that in concrete terms: it can transcribe one hour of audio in under 15 seconds, down from 53 seconds with its predecessor.

MAI-Transcribe-1.5 is now up to 5x more efficient than Gemini 3.1 Flash, Scribe V2, and GPT-4o-Transcribe on the Artificial Analysis leaderboard. The speed gains come specifically from improvements in how the model handles long-form audio, which is where most production workloads actually live.

Where it sits on the leaderboard

MAI-Transcribe-1.5 comes in at 3rd overall on the Artificial Analysis AA-WER leaderboard, behind Alibaba's Fun-Realtime-ASR-preview (1.7% WER) and ElevenLabs Scribe v2 (2.2% WER). But the accuracy story is stronger on the multilingual front. On the FLEURS benchmark across 43 languages, MAI-Transcribe-1.5 achieves an average WER of 4.86%, beating ElevenLabs Scribe v2 (5.53%), OpenAI's transcription model (5.73%), and Google Gemini Flash Lite (5.63%).

Alpha Signal

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves