HeyGen Voice Beats ElevenLabs and Alibaba to Top TTS Rankings

HeyGen's new in-house TTS model debuts on top of the Controlled Voice Arena, beating Qwen-Audio-3.1-TTS-Plus and ElevenLabs' Eleven v4 Turbo.

·
·
·
Read5 min
TypeNews
  • HeyGen Voice debuts at #1 on Artificial Analysis' Controlled Voice Arena with 1,201 Elo.
  • Beats Qwen-Audio-3.1-TTS-Plus (1,182), Eleven v4 Turbo (1,166) and Eleven v4 (1,162) in blind tests.
  • Ranks #1 for Assistants, Customer Service and UK English; #2 for US English behind Qwen.
  • Priced at $30 per 1M characters, between Qwen ($19.3) and Eleven v4 Turbo ($40); runs ~40 chars/sec.
  • Weaker on pronunciation robustness (83.1%, #10 of 29), especially expanding shorthand like abbreviations.
  • Available via HeyGen's v3 API with per-speaker professional voice cloning from 20+ minutes of audio.

HeyGen Voice tops a controlled TTS ranking

HeyGen has released its first in-house standalone text-to-speech model, HeyGen Voice docs, and it leads the referenced snapshot of Artificial Analysis’ Controlled Voice TTS Arena. Blind evaluators preferred it over Alibaba’s Qwen-Audio-3.1-TTS-Plus and ElevenLabs’ Eleven v4 models, giving developers another high-ranking option for assistants, support systems, and avatar pipelines.

Eight voices, one controlled test

Artificial Analysis designed its Controlled Voice Arena to reduce the influence of voice selection. Every model clones the same eight reference voices: two male and two female voices in US English, plus two male and two female voices in UK English. Each reference uses the same one- to two-minute source sample, and every model synthesizes the same scripts.

Blind evaluators compare the resulting clips, and those preferences feed an Elo rating. Elo measures performance relative to the other models and votes in the arena, so narrow score differences are directional signals rather than absolute quality measurements.

A lead built on conversational speech

HeyGen Voice recorded an overall Elo score of 1,201 across 1,468 arena appearances. Qwen-Audio-3.1-TTS-Plus followed at 1,182, with Eleven v4 Turbo at 1,166 and Eleven v4 at 1,162.

HeyGen Voice results by benchmark segment
Segment Elo Placement
Overall 1,201 No. 1
Assistants 1,214 No. 1
Customer service 1,212 No. 1
English, UK 1,190 No. 1
Knowledge sharing 1,158 No. 2
English, US 1,206 No. 2, behind Qwen at 1,220
Entertainment 1,178 No. 4

HeyGen’s strongest results come from practical conversational workloads. It ranks first for assistants and customer service, while its lower entertainment placement suggests that teams producing dramatic or highly expressive speech should run separate listening tests.

Abbreviations expose the weak spot

Artificial Analysis also evaluates pronunciation robustness by asking human reviewers to compare difficult passages with agreed pronunciations. HeyGen Voice scored 83.1% overall, placing No. 10 among 29 models.

Pronunciation robustness results
Category HeyGen Voice Category leader
Preserving exact sequences 83.3%, No. 2 SpaceXAI TTS, 85.7%
Standalone terms 90.5% Eleven v3 Conversational, 95.1%
Contextually appropriate pronunciation 89.8% Gemini 3.8 Flash TTS, 97.9%
Expanding shorthand 77.0% Eleven v4, 94.1%

The largest gap appears in shorthand expansion, which covers context-sensitive forms such as reading “Dr.” as “doctor” or “Rd.” as “road.” Its second-place result for exact sequences makes it more competitive on identifiers such as order numbers and codes.

Abbreviation-heavy applications can reduce pronunciation errors with an upstream normalization layer. Production tests should cover addresses, dates, currencies, product codes, acronyms, and domain-specific terms while preserving strings that must be spoken character by character.

Mid-market price, moderate throughput

HeyGen Voice costs $30 per 1 million characters and reports throughput of roughly 40 characters per second. That places its price between Qwen and the two ElevenLabs models in this comparison.

Published prices for the compared models
Model Price per 1 million characters
Qwen-Audio-3.1-TTS-Plus $19.30
HeyGen Voice $30
Eleven v4 Turbo $40
Eleven v4 $80

Eleven v4 reports 76 characters per second, while Gemini 3.8 Flash TTS reports 46. Those figures describe aggregate generation rate. Voice-agent teams should also measure time to first audio, chunk cadence, concurrency, and tail latency under production traffic.

Calling the v3 API

HeyGen Voice is available through the platform and the v3 API reference. HeyGen trains a dedicated adapter for each professional voice and synthesizes speech from that adapter.

  1. Provide one to 10 recordings of the same speaker, totaling at least 20 minutes.
  2. Start voice training and poll its status until it becomes ACTIVE.
  3. Submit text to the synchronous endpoint or use the streaming endpoint for incremental audio.

Python request example

bash
import requests

response = requests.post(
    "https://api.heygen.com/v3/models/audio/tts",
    headers={
        "x-api-key": "YOUR_KEY",
        "Content-Type": "application/json",
    },
    json={
        "voice_id": "vc_...",
        "text": "Your rental upgrade is approved.",
        "language": "en",
    },
    timeout=(10, 300),
)

response.raise_for_status()
print(response.json())

The synchronous endpoint keeps the request open while synthesis runs, then returns JSON containing a URL for a mono PCM16 WAV file sampled at 44.1 kHz. It accepts an ACTIVE professional voice and allows 30 requests per minute for each workspace member.

The streaming endpoint at /v3/models/audio/tts/stream sends ordered audio parts through Server-Sent Events and can include word timestamps. That interface supports latency-sensitive agents, although production suitability still depends on measured startup time, chunk delivery, and interruption handling.

Voice-cloning workflows also require explicit speaker authorization and documented controls for recordings, generated audio, retention, access, and deletion.

First-party speech changes the stack

HeyGen built its business around AI avatars and historically integrated third-party speech providers, including ElevenLabs. Adding first-party voice generation to the Avatar V stack can reduce integration boundaries for teams already using HeyGen and consolidate avatar and speech usage under one vendor relationship.

Lip sync, identity consistency, API reliability, and end-to-end latency require separate evaluation because the cited TTS benchmarks measure listener preference and pronunciation. Teams considering consolidation should test complete avatar renders rather than infer those qualities from the voice ranking.

Re-run the vendor decision

Earlier snapshots of the same leaderboard placed Cartesia Sonic 3.5 first at 1,122 Elo, followed by ElevenLabs Eleven v3 at 1,088 and Inworld Realtime TTS-2 Research Preview at 1,070. The lead has since rotated among Cartesia, Alibaba, and HeyGen. Changes to the model pool and vote distribution also affect Elo, so the 79-point difference between the earlier and current leaders does not represent a direct measurement of industry-wide improvement.

The current comparison still gives buyers a concrete reason to refresh internal evaluations: the highest-ranked models span several vendors, and published prices range from $19.30 to $80 per million characters, a spread of roughly 4.1 times.

  • Voice fit: Compare authorized speakers on production scripts, accents, pacing, and emotional range.
  • Text robustness: Test identifiers, dates, amounts, abbreviations, names, and specialized terminology.
  • Agent latency: Measure time to first audio, chunk intervals, interruption behavior, and p95 end-to-end latency.
  • Reliability: Exercise rate limits, retries, malformed input, long passages, and failed URL delivery.
  • Total cost: Include retries, voice training, storage, streaming infrastructure, and downstream avatar generation.
Trending
  • No trending articles

Comments

avatar

Next Reads