Gradium's Voice Design Beats ElevenLabs With 72.6% Win Rate for Free

Gradium's new Voice Design turns a written description into a brand-new synthetic voice in seconds, no cloning or licensing required.

·
·
Read4 min
TopicAudio · Api
  • Gradium launched Voice Design, generating new synthetic voices from a text prompt in seconds.
  • Free on every plan, including the free tier, available in Studio and API today.
  • Supports English, French, German, Spanish, Portuguese with 100+ regional accents inside each.
  • Won 72.6% of 7,627 blind pairwise comparisons against ElevenLabs, Inworld, MiniMax, Fish Audio.
  • Voices are fully synthetic, no actor licence, consent, or per-voice royalties required.
  • Returns a stable voice_id that runs on the same streaming TTS endpoint as catalog voices.

Paris-based voice AI startup Gradium has launched Voice Design, a text-prompt-to-voice tool that generates entirely new synthetic voices from a sentence or two of description. Type in an accent, an age band, a pitch, and a use case, and candidate voices come back in seconds, each usable through the same TTS endpoint you already call.

The feature is live in Studio and the API today, free on every plan including the free tier. That pricing decision puts it in direct competition with ElevenLabs, Inworld, MiniMax, and Fish Audio, all of which charge for comparable functionality.

Fewer contracts, more voices

Gradium supports English, French, German, Spanish, and Portuguese, with regional accents within each. Voice cloning requires sourcing a real speaker, getting consent, and signing a licence for every voice you ship. Voice Design skips all of that: the voice is fully synthetic, so there is no actor to pay, no licence to renew, and audio generated with it can be used commercially without per-voice royalties.

For teams building voice agents at scale, the bottleneck has shifted from model quality to procurement and legal. A Québécoise receptionist, a Paulista support agent, and a sixty-year-old narrator used to mean three separate speaker contracts. Voice Design collapses that into three prompts.

Benchmark numbers

Gradium published a detailed benchmark alongside the release, covering 7,627 blind pairwise comparisons across six voice-design systems and five languages, judged by native speakers with system names hidden and a shared line spoken in the target language.

Native speakers picked Gradium voices over ElevenLabs, Inworld, MiniMax, and Fish Audio in 72.6% of accent-prompt comparisons, 13.6 points ahead of the next system and first in every language tested. Win rates by system:

  • Gradium: 72.6%
  • ElevenLabs v3: 59.0%
  • Inworld: 44.8%
  • Fish Audio: 36.7%
  • MiniMax: 31.7%

The largest gaps appeared on regional accents that most TTS systems flatten. Voice Design reached a 97% win rate on Québécois French, 86% on Rioplatense Spanish, and 85% on Bavarian German. Running the same clips through Gemini 3.1 Pro as an LLM judge on a 1-to-5 scale produced the same ranking: Gradium at 4.06, ElevenLabs at 3.86.

How it fits into a codebase

The API returns up to five candidate voices per prompt. Audition them, pick one, name it, and Gradium returns a stable voice_id that drops into the same streaming TTS endpoint with identical latency and output formats. The model samples rather than retrieves, so the same description submitted twice produces fresh voices each time. Save the voice_id after selection or it cannot be regenerated from the prompt alone.

Prompts can specify:

  • Region or diaspora (Colombian rather than Latin American, Québécois rather than French Canadian)
  • Age band, gender, pitch, pace, and energy
  • Vocal texture (vocal fry, warm resonance, sparkling)
  • Intended use case (support agent, narrator, receptionist)

Limitations

Reproducing a specific real person still requires voice cloning. Coverage is limited to five languages, so anything outside English, French, German, Spanish, and Portuguese is not supported. And because sampling is non-deterministic, saving the voice_id on first generation is the only way to return to a chosen voice later.

Company background

Voice Design arrives roughly nine months after Gradium emerged from stealth. The company reopened its seed round to new investors including Nvidia and raised $100 million total, using the capital to open a Bay Area office. Gradium was spun out of French AI lab Kyutai and co-founded by Neil Zeghidour, a researcher who previously worked at Google Brain, DeepMind, and Facebook.

ElevenLabs was valued at $11 billion in February, and every frontier lab now ships some form of voice model. Gradium's bet is that winning on niche accents and pricing custom voices at zero is enough to pull developers away from incumbents, especially in markets where a Parisian accent or a Chilango register is the difference between a product that resonates and one that does not.

Comments

avatar