Sarvam Opens Content Studio to Dub Videos Into 11 Indian Languages
Sarvam opens its multilingual voice and video workspace to the public, bundling dubbing, cloning, and TTS across 11 Indian languages into one browser tool.
- Sarvam Content Studio is now open to all, bundling dubbing, TTS, voice cloning, and STT in one browser workspace.
- Video dubbing preserves the original speaker's voice, tone, and rhythm across 11 Indian languages.
- Users can clone a voice from a short recording and reuse it as a persona.
- Built on Bulbul V3 TTS, Saaras STT, and Sarvam Translation, with 35+ voices available.
- Fine-grained editing: review every line, adjust pronunciation, generate subtitles before export.
- Try it at indus.sarvam.ai/creator-studio.
Sarvam has opened Content Studio, its browser-based workspace for generating, cloning, and localizing audio and video in Indian languages. Feed it a script and get natural-sounding speech, upload a short recording and get a reusable voice clone, or drop in a video and get it dubbed into another language while the original speaker's voice, tone, and rhythm carry over.
These capabilities previously lived behind API access and enterprise pilots. Making the dashboard self-serve turns Sarvam's stack into something a solo creator or a two-person marketing team can use without writing code.
Inside the workspace
Content Studio bundles four core tools that were previously accessible mainly as separate APIs:
- Dubbing: localize video and audio into Indian languages with speaker voice cloning, tone control, and multi-format exports.
- Text to Speech: generate voiceovers from a library of pre-built personas.
- Voice Cloning: map a speaker's vocal traits into a reusable persona from a short sample.
- Voice Library: browse curated voices alongside your own clones in one place.
- Speech to Text: transcribe audio across regional dialects with speaker diarization.
The dubbing pipeline stitches together three of Sarvam's in-house models. Saaras, the speech-to-text model, transcribes the original audio and extracts dialogue with timestamps. Sarvam Translation converts the script while preserving the timing cues needed to keep lip-sync intact. Bulbul V3, the TTS model, generates the voice audio. After generation, users can review every line, tweak translations, adjust pronunciation, and export subtitles before shipping.
Why Indian languages, specifically
Sarvam's argument is that generic Western TTS systems still treat Indian languages as an afterthought, fine-tuning them from English models with limited data. Bulbul V3 was trained from scratch on Indian speech, which matters for two things global models routinely botch: code-switching between English and an Indian language mid-sentence, and correct pronunciation of Indian names, addresses, and abbreviations.
The platform covers Hindi, Tamil, Telugu, Bengali, Marathi, Malayalam, Kannada, Gujarati, Punjabi, Odia, and English, with more than 35 voices across those languages. In a third-party listener study cited by the company, Bulbul V3 produced the highest listener preference and the lowest error rates across use cases and languages.
The economics it rewrites
Traditional dubbing for a single language takes three to six weeks once you account for script translation, artist booking, studio recording, sync, and QA. That timeline breaks for anything published on a weekly cadence, and it collapses entirely for lower-resource languages where the artist pool is small.
Content Studio compresses that loop to minutes. When dubbing costs drop from lakhs to hundreds of rupees, the calculus shifts from which content to dub to why anything is left undubbed. A YouTube creator running a Hindi channel can spin up Tamil, Telugu, and Bengali versions from the same script without hiring parallel studios.
The company behind it
Sarvam is a Bengaluru outfit founded in August 2023 by IIT alumni Vivek Raghavan and Pratyush Kumar, both previously associated with AI4Bharat at IIT Madras. It sits at an unusual intersection of venture-backed startup and national infrastructure project. The Government of India selected Sarvam to build the country's sovereign Large Language Model under the IndiaAI mission. That first sovereign LLM will not be open-sourced, and a government body will take equity in Sarvam in exchange for the investment.
Content Studio is the consumer-facing surface of a broader stack, which now includes the Sarvam-30B and Sarvam-105B foundational models unveiled at Bharat Mandapam during the India AI Impact Summit, plus the Indus consumer app built on top of them.
What to watch for
For anyone shipping content in Indian languages, the interesting question is whether the browser tool serves as a shortcut into the API or a destination in its own right. Sarvam is positioning it as both. Sarvam Content Agents offers a full video dubbing pipeline where you upload a video, translate it, generate the voiceover, and sync the audio, all as a single workflow for multi-language video production. The combination supports end-to-end localization for OTT and corporate content.
The wider bet is that Indian-language content is one of the largest underserved audio markets on the internet, and that owning the full model stack for it, from speech recognition to TTS to voice cloning, is a defensible position that global players fine-tuning English-first models will struggle to match.