

TTS benchmarks have always had a dirty secret: a model that ships beautiful default voices looks great on a leaderboard, even if its underlying synthesis engine is mediocre. Artificial Analysis just closed that loophole with the launch of its Controlled Voice Arena Leaderboard, a new evaluation that clones the exact same 8 voices across every model and then lets human listeners vote blind.
The variable that was skewing every benchmark
The existing Provider Voice Arena lets each model use its own curated voices. That is fine for comparing finished products, but it conflates two very different things: how good a model's synthesis engine is, versus how good its voice designers are. If ElevenLabs ships a perfectly tuned US female voice and a competitor ships a mediocre one, the competitor loses even if its underlying model is technically superior.
The Controlled Voice Arena standardizes the comparison by cloning the same 1-2 minute reference recordings across every model. Each model is evaluated on the same 8 voice categories: 2 US Male, 2 US Female, 2 UK Male, and 2 UK Female voices. The result is a leaderboard that measures synthesis quality in isolation, not voice curation skill.
How the scoring works
The evaluation uses blind human preference testing. Listeners are presented with pairs of speech clips generated from identical prompts, without knowing which provider produced which clip, and simply select the one they prefer. This eliminates brand bias and ensures rankings reflect the actual listening experience. Those preference judgments are aggregated using an Elo rating system, the same framework used in competitive chess and LMSYS Chatbot Arena for evaluating large language models.
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves
