Sarvam AI's Saaras V4 Beats Global English Benchmarks With Single-Pass Multi-Speaker Recognition
Sarvam's new Saaras V4 Multi-speaker handles overlapping conversations in a single pass, while claiming state-of-the-art English ASR across seven global benchmarks.
- Saaras V4 Multi-speaker announced at Epoch 2026: single-pass diarization + ASR that preserves overlapping speech, compared live against ElevenLabs Scribe v2 and Deepgram.
- First India-built ASR to claim state-of-the-art English averaged across 7 global benchmarks, including 6 foreign/global ones.
- Covers 22 Indian languages with 5 output modes: verbatim, transcribe, code-mixed, transliteration, and translation.
- Sarvam open-sources a 108-hour, 22-language, ~1,200-speaker Indic multi-speaker conversational dataset on Hugging Face.
- Available now via Sarvam API at Rs. 1.5/min with free credits on signup.
- Part of a broader Epoch 2026 wave including Sarvam 105B upgrades, Bulbul V4 TTS, Vision 2.0, and a sovereign inference platform.
Sarvam AI used its first Epoch 2026 keynote in Bengaluru to announce Saaras V4 Multi-speaker, a speech recognition model built to handle the hardest audio scenario in real-world deployments: multiple people talking at the same time. It also marks the first time an India-built ASR system claims state-of-the-art performance on global English benchmarks.
One pass instead of three
The conventional pipeline for multi-speaker audio runs three separate steps: a diarizer figures out who is speaking when, cuts the audio into chunks, and then an ASR model transcribes each chunk sequentially. Overlapping speech gets dropped entirely. Saaras V4 Multi-speaker collapses this into a single forward pass using a speaker-biased architecture that performs diarization and transcription simultaneously, preserving overlaps and back-channels that the old pipeline would discard.
At the keynote, Sarvam demoed the model separating an overlapping male speaker and female anchor, comparing output directly against ElevenLabs Scribe v2 and Deepgram Diarization v2. The unified architecture also handles Indic languages, English, and code-mixed speech without switching between models.
English SOTA, claimed for the first time from India
Saaras V4 was evaluated on Svarah, a 9.6-hour benchmark covering 117 speakers across 65 districts in 19 Indian states, with both read and spontaneous conversational speech. The bigger claim is on global English: averaged across seven standard English benchmarks, six of them international, Saaras V4 achieves state-of-the-art performance. No India-built model has made that claim before.