Sarvam AI's Saaras V4 Beats Global English Benchmarks With Single-Pass Multi-Speaker Recognition

Sarvam's new Saaras V4 Multi-speaker handles overlapping conversations in a single pass, while claiming state-of-the-art English ASR across seven global benchmarks.

·
·
AuthorSarvam
Read2 min
  • Saaras V4 Multi-speaker announced at Epoch 2026: single-pass diarization + ASR that preserves overlapping speech, compared live against ElevenLabs Scribe v2 and Deepgram.
  • First India-built ASR to claim state-of-the-art English averaged across 7 global benchmarks, including 6 foreign/global ones.
  • Covers 22 Indian languages with 5 output modes: verbatim, transcribe, code-mixed, transliteration, and translation.
  • Sarvam open-sources a 108-hour, 22-language, ~1,200-speaker Indic multi-speaker conversational dataset on Hugging Face.
  • Available now via Sarvam API at Rs. 1.5/min with free credits on signup.
  • Part of a broader Epoch 2026 wave including Sarvam 105B upgrades, Bulbul V4 TTS, Vision 2.0, and a sovereign inference platform.

Sarvam AI used its first Epoch 2026 keynote in Bengaluru to announce Saaras V4 Multi-speaker, a speech recognition model built to handle the messiest audio scenario in real-world deployments: multiple people talking at the same time. The release also marks the first time an India-built ASR system claims state-of-the-art performance on global English benchmarks.

One pass instead of three

The core problem Saaras V4 Multi-speaker solves is architectural. Saaras V4 handles Indic languages, English, code-mixed speech, and multi-speaker audio in a unified model. The conventional pipeline for multi-speaker audio runs three separate passes: a diarizer (which figures out who is speaking when) cuts the audio into chunks, an ASR model transcribes each chunk one at a time, and overlapping speech gets dropped entirely in the process. Saaras V4 Multi-speaker collapses this into a single forward pass using a speaker-biased model that does diarization and transcription simultaneously, preserving overlaps and back-channels that the old pipeline would simply lose.

On stage, Sarvam demoed the model separating an overlapping male speaker and female anchor, comparing it directly against ElevenLabs Scribe v2 and Deepgram Diarization v2. Saaras V4 Multi-speaker can separate and transcribe multiple people talking at the same time, making meeting recordings much easier to understand.

English SOTA, claimed for the first time from India

Saaras V4 was evaluated on Svarah, a 9.6-hour benchmark comprising 117 speakers across 65 districts in 19 Indian states, capturing substantial accent variation and both read and spontaneous conversational speech. But the bigger claim is on global English: Saaras V4 achieves state-of-the-art performance averaged across seven standard English benchmarks, six of which are foreign and global , the first time an India-built model has made that claim.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves