Cartesia's Sonic 3.6 Dominates Eight of Nine Global Voice AI Leaderboards

Artificial Analysis extended its Controlled Voice Arena to nine new languages, revealing sharp performance gaps between models across Hindi, Mandarin, Arabic and more.

·
·
Cartesia's Sonic 3.6 Dominates Eight of Nine Global Voice AI Leaderboards
Read4 min
TypeNews
  • Artificial Analysis launched Controlled Voice TTS leaderboards for 9 new languages beyond English.
  • Cartesia Sonic 3.6 leads 7 of 9 languages; Sonic 3.5 wins Portuguese.
  • Inworld Realtime TTS-2 wins Mandarin at 1,185 Elo, beating Sonic 3.6.
  • ElevenLabs Eleven v3 family places in top 3 across 7 of 9 languages.
  • Over 100,000 human preference votes collected, with 15 to 24 models ranked per language.
  • Open weights leaders vary by language: Higgs Audio V3, OpenAudio S1 Mini, Voxtral, Breeze TTS 2.

TTS leaderboards reshuffle across nine languages

Artificial Analysis has expanded its Controlled Voice Arena beyond English, adding leaderboards for Spanish, French, German, Portuguese, Japanese, Hindi, Mandarin Chinese, Arabic, and Vietnamese. More than 100,000 native-speaker preference votes rank 15 to 24 public models in each language.

English-first benchmarks can hide substantial regional differences. Cartesia’s Sonic family leads eight of the nine new boards, but the winning version varies in Portuguese, while Inworld takes first place in Mandarin. The results give developers a stronger basis for choosing models by target market instead of applying an English ranking globally.

Cloned voices narrow the comparison

The arena reduces the influence of default voice identity by making competing models reproduce shared reference voices. Each language uses one professionally recorded male voice and one professionally recorded female voice, both from native speakers.

Listeners can therefore focus more closely on naturalness, pronunciation, pacing, prosody, audio quality, and cloning fidelity. Voice cloning cannot remove every model-specific variable, but it limits the advantage a vendor might gain from supplying a particularly appealing default voice.

Artificial Analysis screens raters as first-language speakers and recruits people living in their country of origin where possible. A model appears only on leaderboards for languages it officially supports.

The site derives arena-style Elo scores through a linear-regression method similar to the approach used for LMSYS Chatbot Arena. Higher scores indicate stronger listener preference within that language’s model pool. Because every language has a separate set of voters and models, ranks and score gaps are most useful within each leaderboard rather than as a shared scale across languages.

Cartesia takes eight boards

Cartesia’s Sonic 3.6 ranks first in Spanish, French, German, Japanese, Hindi, Arabic, and Vietnamese. Sonic 3.5 wins Portuguese, while Inworld’s Realtime TTS-2 leads Mandarin.

Language Leading results
Spanish Sonic 3.6 ranks first.
French Sonic 3.6 scores 1,349 Elo, 71 points ahead of Eleven v3 Conversational at 1,278.
German Sonic 3.6 leads Inworld Realtime TTS-2 by 22 points. Inworld scores 1,261.
Japanese Sonic 3.6 ranks first, followed by Inworld Realtime TTS-2 at 1,123.
Hindi Sonic 3.6 ranks first.
Arabic Sonic 3.6 scores 1,369, leading Eleven v3 by 132 points.
Vietnamese Sonic 3.6 scores 1,721, followed by Eleven v3 Conversational at 1,677 and Eleven v3 at 1,676.
Portuguese Sonic 3.5 scores 1,324, ahead of Sonic 3.6 at 1,291.
Mandarin Inworld Realtime TTS-2 scores 1,185, ahead of Sonic 3.6 at 1,146 and StepFun StepAudio 2.5 TTS at 1,130.

ElevenLabs supplies the broadest second tier. Eleven v3 and v3 Conversational place in the top three across seven of the nine languages, including a near tie for second place in Vietnamese.

Inworld’s Realtime TTS-2 is the most consistent challenger outside ElevenLabs. Along with its Mandarin win, it places second in Japanese and German and fourth in French and Portuguese. It also leads the English leaderboard at 1,144 Elo, one point above Sonic 3.6. The listed prices are $20.80 per million characters for Inworld and $49 for Cartesia, though current pricing should be verified before deployment.

Open weights split four ways

Teams that require downloadable weights face a fragmented field. Four model families lead the open-weight entries across the nine languages:

Each language includes only two to six open-weight entries, compared with 15 to 24 models overall. Some markets show gaps of several hundred Elo points between the strongest open-weight and proprietary systems. Those gaps measure listener preference under the arena’s conditions; deployment decisions also depend on licensing, hardware requirements, inference speed, quantization support, and operational cost.

Turn the rankings into a test plan

The rankings support a language-by-language selection process. A Mandarin team would shortlist Inworld, Sonic 3.6, and StepFun in that order, while a Portuguese team should include Sonic 3.5 even though Sonic 3.6 is newer. Small Elo differences offer limited grounds for a final choice, especially when models sit only a few points apart.

  1. Shortlist the leading models on every target-language board.
  2. Test product-specific prompts containing names, numbers, abbreviations, dates, addresses, and code-switching.
  3. Recruit listeners from the intended market and include relevant regional accents or dialects.
  4. Measure time to first audio, streaming stability, concurrency limits, uptime, and total serving cost.
  5. Verify voice-cloning consent requirements, commercial licenses, data residency, and retention policies.

The arena covers perceived output quality under controlled voice conditions. Production testing must also account for API reliability, latency, customization, safety controls, and how well each model handles the application’s actual text distribution.

Developers can inspect the full Controlled Voice leaderboards, submit comparisons through the public arena, and review the methodology notes, including the separate pronunciation-robustness tests.

Trending
  • No trending articles

Comments

avatar

Next Reads