Sarvam AI Cuts Transcription Errors 20% for India's Wealth Manager Dezerv

Dezerv built a post-call analytics system on Sarvam's Speech to Text API, cutting transcription error rates by 20% across 3.6 lakh minutes of client calls monthly.

·
·
Sarvam AI Cuts Transcription Errors 20% for India's Wealth Manager Dezerv
AuthorSarvam
Read2 min
  • Sarvam AI and Dezerv published a case study on a post-call analytics system built on Sarvam's Speech to Text API.
  • Dezerv now processes 3.6 lakh minutes (~60,000 hours) of client calls monthly through Sarvam's API.
  • Word Error Rate dropped ~20% vs. Dezerv's previous outsourced transcription pipeline.
  • The system handles code-mixed Indian speech (Hinglish, regional languages) with built-in speaker diarization across video and phone calls.
  • Dezerv manages over ₹17,000 crore in AUM and operates across five Indian cities with 650+ staff.
  • Zero-Retention AI Processing ensures no client-identifying data is stored with Sarvam, meeting SEBI-regulated compliance requirements.

India's wealth management firm Dezerv has published a case study with Sarvam AI detailing how it built a post-call analytics system on Sarvam's Speech to Text API. The numbers make a strong argument for India-first speech models in high-stakes financial contexts.

The insight locked inside every client call

Dezerv's business runs on relationship managers. Every client is assigned one, and RMs explain portfolio decisions, steady clients through volatile markets, and carry the conversations on which long-term trust is built. Those conversations happen over video conferencing, in whichever language the client is most comfortable using.

Dezerv's RM calls were the company's richest source of customer insight and its least accessible. Every call revealed what clients worried about, which explanations landed, and where RMs needed coaching. With thousands of calls every month, all of that was locked inside audio recordings nobody could search.

The previous transcription pipeline compounded the problem. It was outsourced end to end, so Dezerv could not audit its own quality. When the team ran comparisons, the transcripts fell short, especially for Indic languages.

Why generic ASR breaks on Indian speech

The transcription challenge is real. Dezerv's clients and RMs speak the way urban India speaks: mixing English, Hindi, and regional languages within a single sentence, with dense references to fund names and numbers, over video conferencing calls. This code-mixing, switching languages mid-sentence, is something most Western-trained ASR (automatic speech recognition) models handle poorly.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves