Sarvam AI Ships Bulbul V4 to Bring Emotional Voice to 11 Indian Languages
Sarvam AI's Bulbul V4 upgrades its Indic TTS engine with richer emotion, natural expression, and a wider vocal range across 11 Indian languages.
- Bulbul V4 launched: Sarvam AI released its newest Indic TTS model with richer emotion, natural expression, and greater vocal range.
- Builds on V3's benchmark wins: V3 already beat ElevenLabs and Cartesia Sonic-3 in a 20,000-vote blind study across 11 Indian languages.
- LLM-based prosody engine: The model uses an LLM to infer emphasis, pauses, tone, and pacing from context rather than treating text as a flat sequence.
- 11 languages, 39+ voices, sub-250ms streaming: Supports Hindi, Tamil, Telugu, Bengali, and 7 more; REST and WebSocket APIs available.
- Pricing: Free tier with ₹1,000 credits; paid tier at ₹30 per 10,000 characters; enterprise discounts available.
- Key use cases: Voice agents, BFSI call centers, EdTech tutors, audiobooks, gaming characters, and IVR systems across India.
Sarvam AI just shipped Bulbul V4, the latest version of its text-to-speech model built specifically for Indian languages. The improvements center on richer emotion, more natural expression, and a wider vocal range, building directly on a V3 architecture that already beat ElevenLabs and Cartesia in independent blind listening tests.
Why Indian TTS is a different problem
Building TTS for India means handling speech that switches languages mid-sentence, varies by region, and mixes scripts. Bulbul handles code-mixed text like Hinglish, number normalization, and natural prosody out of the box. The global model providers do English well and a handful of European languages adequately. For a farmer in Maharashtra asking a government AI about crop subsidies in Marathi, or a patient in Tamil Nadu trying to understand a prescription in Tamil, those models fall short. Bulbul V4 is Sarvam's answer to that gap.
What V3 established
Bulbul V3 uses an LLM to analyze text and infer prosodic elements: emphasis, pauses, tone, and pacing. Prosody is the rhythm, stress, and intonation that makes a question sound like a question rather than a statement. By understanding context and intent rather than treating words as a flat sequence, V3 generates speech that aligns with the emotional content of what is being said.