Hume AI's VoiceEQ Benchmark Exposes What 1M Human Ratings Reveal About Voice Models
Hume AI's new benchmark evaluates 40+ voice models on human-quality dimensions that word error rate and latency completely miss
- Hume AI launched Real World VoiceEQ, a benchmark evaluating 40+ voice models on human-quality dimensions beyond WER and latency.
- Built from over 1 million human ratings (785K TTS, 48K S2S), making it one of the largest human voice AI evaluations ever.
- Key finding: no single model dominates -- voice AI is specializing, and no TTS system ranked top-5 across all 8 capability groups.
- S2S models showed the widest performance gap: some recognize emotion well but fail to respond naturally; many remain transcript-driven despite having audio access.
- Automated LLM judges (SLMs) fall short for subjective voice tasks; human raters remain essential for emotional fit and identity consistency.
- The public leaderboard is free on Hugging Face; custom private evaluations are available via Hume's Kairos platform.
Voice AI benchmarks have had a dirty secret for a while: they measure the wrong things. Word error rate (WER) tells you if a model transcribed correctly. Latency tells you if it responded fast. Neither tells you whether a hesitant "...yeah" was understood differently from a confident "yeah" -- and in a fraud alert or a medical conversation, that distinction is everything. Hume AI is now putting a number on that gap.
Hume has launched Real World VoiceEQ, a public benchmark that evaluates voice AI on the dimensions humans actually care about: tone, pacing, hesitation, emotional fit, speaker consistency, and conversational appropriateness. It is free to explore on Hugging Face right now.
The biggest dataset of human voice ratings ever assembled
Real World VoiceEQ evaluates more than 40 leading proprietary and open-source voice models across 15+ key evaluation dimensions and more than 60 metrics, spanning ASR, TTS, Speech-to-Speech (S2S), and Speech Understanding. The scale of the human data behind it is what makes this unusual.
The benchmark was developed from more than 1 million individual human ratings collected across different demographics, speaking styles, and acoustic environments. The current release includes 785,000 TTS ratings and 48,000 STS ratings, making it one of the largest human evaluations of voice AI conducted to date.
Every evaluation was conducted using Kairos, Hume's flexible, voice-native evaluation platform. The same infrastructure enables frontier AI labs and enterprises to run custom evaluations tailored to specific use cases, identify granular failure modes in production voice systems, generate human preference data, and continuously improve models through reinforcement learning and human feedback.
What the leaderboard actually measures
The benchmark is split across four major task categories:
- Text-to-Speech (TTS): naturalness, expressiveness, emotional delivery, voice consistency, and instruction-following (e.g., "whisper like you're sharing a secret")