Alibaba's Qwen-Audio-3.0-TTS Tops Speech Arena With 16-Language Voice Cloning

Alibaba's Qwen-Audio-3.0-TTS tops the global TTS leaderboard with 16-language support, inline emotion tags, and noise-robust voice cloning

·
·
Alibaba's Qwen-Audio-3.0-TTS Tops Speech Arena With 16-Language Voice Cloning
AuthorTongyi Lab
Read2 min
TopicAudio · Api
  • Leaderboard #1: Qwen-Audio-3.0-TTS-Plus tops the Artificial Analysis Speech Arena with an Elo score of 1,237, beating Simba 3.2, Gemini 3.1 Flash TTS, and Sonic 3.5.
  • Two variants: Flash (real-time, low latency) and Plus (high quality), both available now via Alibaba Cloud Model Studio API.
  • 16 languages + 20 Chinese dialects: Adds Arabic, Indonesian, Portuguese, Thai, Vietnamese, Malay, and Filipino over the previous version.
  • 86 inline control tags: Embed emotions, laughter, sighs, and style shifts directly in text; also accepts free-form natural language instructions.
  • Noise-robust voice cloning: Clones voices from noisy or reverberant reference audio without a separate denoising step; supports up to 3-minute one-pass synthesis.
  • Speed tradeoff: Plus generates only 16 characters/second vs. 120 for Sonic 3.5; priced at $27.59 per 1M characters.

Qwen-Audio-3.0-TTS is Alibaba's latest production-grade text-to-speech system, and it just landed at the top of the Artificial Analysis Speech Arena Leaderboard. It ships in two variants: Flash for real-time, low-latency applications and Plus for maximum output quality. Both are available via API on Alibaba Cloud Model Studio right now.

First place, by a hair

Qwen-Audio-3.0-TTS-Plus currently leads the Text-to-Speech Arena with an Elo score of 1,237. That puts it narrowly ahead of Simba 3.2 at 1,234, and ahead of Gemini 3.1 Flash TTS at 1,214 and Sonic 3.5 at 1,207. The confidence intervals overlap between the top two, so the gap is tight, but the ranking is real. Rankings are based on blind user votes in the Speech Arena, where users listen to pairs of speech samples generated from the same text and choose which sounds more natural. Higher Elo scores indicate a model produces speech preferred more often by listeners.

Qwen-Audio-3.0-TTS-Plus via Alibaba Cloud, Simba 3.2 via SpeechifyAI, and Gemini 2.5 Flash Lite TTS via Google currently offer some of the strongest quality-for-price tradeoffs in the comparison, sitting on the current quality-versus-price frontier.

Leaderboard table comparing TTS models by Elo rating, pricing, and number of voices

What's actually new

The headline features break down into four areas worth understanding individually:

  • 16 languages, 20 Chinese dialects. The model adds 7 new languages over its predecessor, including Arabic, Indonesian, Portuguese, Thai, Vietnamese, Malay, and Filipino. It also handles zero-shot voice cloning across all 20 Chinese regional dialects, from Cantonese to Sichuan to Shanghai dialect.
  • Free-style natural language instructions. Instead of configuring audio parameters, you describe the voice in plain text. Instructions like "slow, calming voice" or "female morning radio host, lively but restrained, fast pace" are interpreted directly. You can also pass structured JSON like

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves