Sarvam AI Ships Bulbul V4 to Bring Emotional Voice to 11 Indian Languages

Sarvam AI's Bulbul V4 upgrades its Indic TTS engine with richer emotion, natural expression, and a wider vocal range across 11 Indian languages.

·
·
AuthorSarvam
Read2 min
  • Bulbul V4 launched: Sarvam AI released its newest Indic TTS model with richer emotion, natural expression, and greater vocal range.
  • Builds on V3's benchmark wins: V3 already beat ElevenLabs and Cartesia Sonic-3 in a 20,000-vote blind study across 11 Indian languages.
  • LLM-based prosody engine: The model uses an LLM to infer emphasis, pauses, tone, and pacing from context rather than treating text as a flat sequence.
  • 11 languages, 39+ voices, sub-250ms streaming: Supports Hindi, Tamil, Telugu, Bengali, and 7 more; REST and WebSocket APIs available.
  • Pricing: Free tier with ₹1,000 credits; paid tier at ₹30 per 10,000 characters; enterprise discounts available.
  • Key use cases: Voice agents, BFSI call centers, EdTech tutors, audiobooks, gaming characters, and IVR systems across India.

Sarvam AI just shipped Bulbul V4, the latest version of its text-to-speech model built specifically for Indian languages. The improvements center on richer emotion, more natural expression, and a wider vocal range, building directly on a V3 architecture that already beat ElevenLabs and Cartesia in independent blind listening tests.

Why Indian TTS is a different problem

Building TTS for India means handling speech that switches languages mid-sentence, varies by region, and mixes scripts. Bulbul handles code-mixed text like Hinglish, number normalization, and natural prosody out of the box. The global model providers do English well and a handful of European languages adequately. For a farmer in Maharashtra asking a government AI about crop subsidies in Marathi, or a patient in Tamil Nadu trying to understand a prescription in Tamil, those models fall short. Bulbul V4 is Sarvam's answer to that gap.

What V3 established

Bulbul V3 uses an LLM to analyze text and infer prosodic elements: emphasis, pauses, tone, and pacing. Prosody is the rhythm, stress, and intonation that makes a question sound like a question rather than a statement. By understanding context and intent rather than treating words as a flat sequence, V3 generates speech that aligns with the emotional content of what is being said.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves