ElevenLabs' Eleven v4 Tops English Voice AI With 91.7% Pronunciation Score

Eleven v4 tops the Provider Voice arena with a 1,319 Elo and hits 91.7% pronunciation accuracy, though it costs nearly 5x Google's Gemini Flash TTS.

·
·
Read6 min
TypeNews
  • Eleven v4 tops the Artificial Analysis Provider Voice arena at 1,319 Elo, beating Cartesia Sonic 3.6 and Gemini 3.8 Flash TTS.
  • Scores 91.7% on pronunciation robustness, the highest ever measured on the benchmark, up from 85.6% for v3.
  • Supports 90+ languages versus 70+ in v3, and adds IPA phonetic control for exact pronunciation.
  • Generates 73.4 characters per second but costs $80 per 1M characters, roughly 5x Google's Gemini 3.8 Flash TTS.
  • New architecture enables voice cloning from a 10-second sample, with Turbo variant hitting 150ms time to first speech.
  • ElevenLabs ARR has climbed to $600M+ following its $500M Sequoia round at an $11B valuation.

Eleven v4 takes the English TTS lead at a premium

At publication, Artificial Analysis’ Provider Voice Arena lists Eleven v4 first among English text-to-speech models. It ranks ahead of Cartesia’s Sonic 3.6 and Google’s Gemini 3.8 Flash TTS while setting the benchmark’s highest composite score for pronunciation robustness.

ElevenLabs launched two variants for different workloads. The flagship v4 targets narration and character performances, while Turbo prioritizes low-latency voice agents. TechCrunch reports that both add expression controls and support more than 90 languages, up from roughly 70 in v3.

Native voices put v4 ahead

The Provider Voice Arena evaluates each model with its own voices and default configuration. Its Elo score summarizes head-to-head listener preferences, with a higher score indicating that evaluators preferred the model more often.

English Provider Voice Arena standings at publication
Rank Model Elo score
1 Eleven v4 1,319
2 Cartesia Sonic 3.6 1,276
3 Gemini 3.8 Flash TTS 1,267

Eleven v4’s score draws on 1,674 evaluation samples. The model also ranks first in each of the arena’s four content categories: customer service, assistants, knowledge sharing, and entertainment.

The Controlled Voice board uses the same cloned voice across systems, reducing the influence of each provider’s voice catalog. Eleven v4 ranks second there with 1,157 Elo, behind Alibaba’s Qwen-Audio-3.1-TTS-Plus at 1,178. The result suggests that ElevenLabs’ native voices and default configuration contribute to its lead on the main board.

Elo scores are relative rankings rather than quality percentages, and they can move as evaluators add votes. A 21-point gap on the Controlled Voice board should therefore inform direct testing, not replace it.

Pronunciation reaches 91.7%

Speech models frequently misread abbreviations, technical terms, context-dependent words, and character sequences such as order numbers. Artificial Analysis tests those cases by asking human reviewers to compare generated clips with agreed pronunciations.

Eleven v4 pronunciation robustness results
Category What it tests v4 result Comparison
Expanding Shorthand Expanding abbreviations correctly 94.1% v3 scored 82.7%
Contextually Appropriate Selecting pronunciation from sentence context 94.1% Gemini 3.8 Flash leads at 97.9%
Standalone Terms Reading terms without sentence context 93.2% v3 Conversational leads at 95.1%
Preserving Exact Sequences Reading identifiers and character strings accurately 78.8% v3 scored 71.2%; SpaceXAI TTS leads at 85.7%

The 91.7% composite is the highest Artificial Analysis has measured for a TTS model. Eleven v4 is also the only model above 93% in three of the four categories.

Exact character sequences remain its weakest category at 78.8%. Applications that read confirmation numbers, medical terms, account identifiers, or product SKUs still need production testing and a fallback for high-risk strings.

Throughput rises alongside price

Artificial Analysis measured Eleven v4 at 73.4 generated characters per second, up from 42.5 for v3. That figure measures synthesis throughput; it does not include network delay, language-model response time, or client-side audio playback.

Listed standard API prices
Model Price per million characters Cost relative to v4
Eleven v4 $80 100%
Cartesia Sonic 3.6 $49 61%
Gemini 3.8 Flash TTS $16.49 21%

Eleven v4 costs about 1.6 times as much as Sonic 3.6 and 4.9 times as much as Gemini 3.8 Flash TTS at listed standard rates. Pricing may vary by plan, volume agreement, and included credits.

A two-week launch promotion reduces v4 to $22 per million characters and Turbo to $11 per million. Eligible Creator+ subscriptions can also use v4 within their monthly credit limits without an additional model surcharge. The standard $80 rate becomes the relevant long-term figure after the promotion expires.

Cloning changes require migration

ElevenLabs rebuilt its voice-cloning stack for v4. Instant cloning can create a voice from a 10-second recording, while professional cloning returns after being unavailable in v3. The company says v4 instant clones outperform professional clones produced with Multilingual v2, although the public arena results do not independently test that specific claim.

Professional and instant clones created before v4 require retraining for effective use with the new model. Applications already calling ElevenLabs’ conversion API can select v4 by changing the model identifier:

php
const audio = await elevenlabs.textToSpeech.convert(
  's3TPKV1kjDlVtZbl4Ksh',
  {
    text: 'The first move is what sets everything in motion.',
    modelId: 'eleven_v4'
  }
);

Clone retraining remains a separate migration step from the SDK change. Teams should compare old and retrained voices for identity, pacing, pronunciation, and consistency before moving production traffic.

Turbo reports a median time to first speech of 150 milliseconds and supports bidirectional streaming. Time to first speech measures the interval between a synthesis request and the first returned audio chunk. End-to-end agent latency also depends on network conditions, language-model generation, buffering, and playback.

The English lead has limits

ElevenLabs’ benchmark lead applies specifically to the English Provider Voice Arena. Cartesia’s Sonic 3.6 still leads eight of the nine non-English language boards tracked by Artificial Analysis, despite v4’s support for more than 90 languages.

Competition also includes Deepgram, Fish Audio, Boson, WellSaid Labs, Google, OpenAI, and self-hosted models. On the Controlled Voice board, Qwen-Audio-3.1-TTS-Plus retains the lead, giving teams with a fixed cloned voice another strong candidate to evaluate.

ElevenLabs reports that enterprise customers generate more than 55% of its revenue. Its annualized recurring revenue has exceeded $600 million, up from roughly $330 million at the start of 2026. The company also raised $500 million in a Sequoia-led round at an $11 billion valuation.

Match the model to the workload

  • English narration and character work: Eleven v4 leads the native-voice preference board and offers expanded expression controls, with a standard price of $80 per million characters.
  • Real-time voice agents: Turbo targets faster conversational loops through streaming and a reported 150-millisecond median time to first speech.
  • Fixed cloned voices: Qwen-Audio-3.1-TTS-Plus leads the Controlled Voice board, making same-voice comparisons especially relevant.
  • High-volume, cost-sensitive synthesis: Gemini 3.8 Flash TTS carries a substantially lower listed character price.
  • Non-English speech: Sonic 3.6 leads most of the language-specific boards tracked by Artificial Analysis.
  • Self-hosted deployment: Open-weight systems such as Breeze TTS 2 exchange per-character API fees for infrastructure and operational work.

Production evaluations should use representative scripts, voices, languages, and identifiers from the intended application. Eleven v4’s leaderboard position makes it a leading candidate for English workloads, while its price, clone-migration requirement, and weaker exact-sequence score remain material deployment considerations.

Trending
  • No trending articles

Comments

avatar

Next Reads