Sarvam AI Ships Bulbul V4 to Bring Emotional Voice to 11 Indian Languages

Sarvam AI's Bulbul V4 upgrades its Indic TTS engine with richer emotion, natural expression, and a wider vocal range across 11 Indian languages.

·
·
AuthorSarvam
Read2 min
  • Bulbul V4 launched: Sarvam AI released its newest Indic TTS model with richer emotion, natural expression, and greater vocal range.
  • Builds on V3's benchmark wins: V3 already beat ElevenLabs and Cartesia Sonic-3 in a 20,000-vote blind study across 11 Indian languages.
  • LLM-based prosody engine: The model uses an LLM to infer emphasis, pauses, tone, and pacing from context rather than treating text as a flat sequence.
  • 11 languages, 39+ voices, sub-250ms streaming: Supports Hindi, Tamil, Telugu, Bengali, and 7 more; REST and WebSocket APIs available.
  • Pricing: Free tier with ₹1,000 credits; paid tier at ₹30 per 10,000 characters; enterprise discounts available.
  • Key use cases: Voice agents, BFSI call centers, EdTech tutors, audiobooks, gaming characters, and IVR systems across India.

Sarvam AI just shipped Bulbul V4, the latest version of its text-to-speech model built specifically for Indian languages. The headline improvements are richer emotion, more natural expression, and a greater vocal range -- a direct evolution from the V3 architecture that already beat ElevenLabs and Cartesia in independent blind listening tests.

The problem Bulbul is solving

Building TTS for India is not the same as building TTS for English and then translating. Indian speech is complex by default -- people switch languages mid-sentence, and accents vary by region. On top of that, Bulbul is purpose-built for Indian languages and accents, handling code-mixed text (like Hinglish), number normalization, and natural prosody out of the box.

The global LLM giants -- OpenAI, Google, Anthropic, Meta -- are excellent at English and adequate at a handful of European languages. For a farmer in Maharashtra asking a government AI about crop subsidies in Marathi, or a patient in Tamil Nadu trying to understand their prescription in Tamil, these models are essentially useless. Bulbul V4 is Sarvam's answer to that gap.

What V3 established (and V4 builds on)

Bulbul V3 is built on an LLM to analyze text and infer the prosodic elements of natural speech -- emphasis, pauses, tone, and pacing. By understanding context and intent rather than processing words as a simple sequence, it generates speech that sounds natural and aligns with the emotional content of what is being said. Prosody here means the rhythm, stress, and intonation that makes a question sound like a question -- the thing that separates natural speech from robotic output.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves