Artificial Analysis' Music Arena v1.1 Crowns Suno v6 Across 1,000 Harder Prompts
Artificial Analysis rebuilt its Music Arena with 1,000 new multi-genre prompts and 37,000 blind panel votes, shaking up how text-to-song models are ranked.
- AA-Music v1.1 launches with 1,000 new prompts across 17 genres and 37,000 blind panel votes.
- Suno v6 ranks #1 on both vocal (Elo 1142) and instrumental (Elo 1140) leaderboards.
- Suno v6-mini takes #2 on both boards, giving Suno the top two slots everywhere.
- Mureka V9.5 reaches #3 on vocals with a 51-Elo jump over Mureka V9.
- MiniMax Music 3.0 is the top open-weights model, anchored at Elo 1000.
- Public arena votes no longer count; only recruited blind panel ratings feed Elo.
Music Arena v1.1 adds 1,000 prompts and panel-only scoring
Artificial Analysis has released v1.1 of its two Music Arena benchmarks, introducing a more demanding prompt set and restricting ranked votes to recruited evaluators. The changes aim to distinguish leading text-to-music systems such as Suno v6, Lyria 3.5 and Mureka V9.5, whose earlier scores clustered near the top.
Detailed briefs test composition, performance and mix
The Music Arena now uses 1,000 newly written prompts, divided evenly between vocal and instrumental tests. The set covers 17 genres, from classical, jazz and blues to EDM, metal and reggae. Prompts specify song structure, mood, instrumentation, production and vocal style.
One sample vocal prompt requests tech house with a sassy, half-spoken female vocal, rubbery wobble bass, swung hi-hats and a dry club mix. Fulfilling that brief requires the model to coordinate genre conventions, arrangement, performance and mix character in one generation.
Ranked Elo scores now use only blind preferences from a recruited evaluator panel. Public arena votes no longer affect the leaderboard, limiting the influence of brigading and casual voting. The launch snapshot includes more than 37,000 blind votes, with about 16,000 for vocals and 21,000 for instrumentals. Every ranked model has received more than 2,000 head-to-head comparisons.
Suno takes both top slots
At launch, Suno v6 ranks first on both boards, followed by Suno v6-mini. Mureka V9.5 places third on the vocal standings, while four models share a statistically tight cluster below Suno on the instrumental benchmark.
| Board | Rank | Model | Elo |
|---|---|---|---|
| Vocal | 1 | Suno v6 | 1142 |
| Vocal | 2 | Suno v6-mini | 1116 |
| Vocal | 3 | Mureka V9.5 | 1103 |
| Instrumental | 1 | Suno v6 | 1140 |
| Instrumental | 2 | Suno v6-mini | 1109 |
| Instrumental | 3 | Mureka V9 | 1081 |
Mureka V9.5 gains 51 vocal Elo points over Mureka V9, the largest generation-to-generation improvement reported for that family. On the instrumental board, Mureka V9 at 1081, Mureka V9.5 at 1080, Lyria 3 Pro at 1078 and Lyria 3.5 at 1076 are statistically tied from third through sixth, spanning only five Elo points.
Signals below the leaders
- End-to-end vocals: The benchmark supplies no lyrics, so each model must write and perform them.
- Open weights: MiniMax Music 3.0 is the highest-ranked open-weights entry, placing No. 11 for vocals and No. 13 for instrumentals. It serves as the 1000-point Elo anchor on both boards.
- Close instrumental challenger: Stable Audio 3 Medium scores 998, within three points of MiniMax Music 3.0.
- Earlier models: Google’s Lyria 2 scores 932, below Lyria 3 Pro and Lyria 3.5. Meta’s MusicGen records 881, the lowest reported score.
One vote captures several qualities
Each comparison presents two tracks generated from the same prompt, and a panelist chooses a favorite without seeing the model names. Elo converts those wins and losses into relative ratings, with frequent winners receiving higher scores. An Elo value is neither a quality percentage nor an absolute measurement, and comparisons across benchmark versions or between the two boards have limited meaning.
Because a panelist makes one preference decision, the score combines prompt adherence, composition, performance, sound quality and personal taste. The benchmark does not assign separate grades to those dimensions, so a track can win through strong production even when another follows the brief more precisely.
Separate vocal and instrumental boards expose different capabilities. Instrumental generation emphasizes composition, arrangement and production, while the vocal task also requires lyric writing, pronunciation, singing style and placement of the voice in the mix. Those added demands explain why some models change position between the boards.
Build a product test from the shortlist
Product teams building game soundtracks, advertising music, podcast beds or user-generated audio can use the boards to narrow their candidate set. A shipping decision still requires workload-specific testing because Elo omits latency, price, throughput, maximum duration, editing controls, stem export, consistency, safety policies and commercial terms.
- Match the workload. Build an internal prompt set that reflects the genres, durations, vocal styles and production constraints your application will use.
- Filter by deployment needs. Suno v6 leads this evaluation as a proprietary model. Teams evaluating open weights can begin with MiniMax Music 3.0 and include Stable Audio 3 Medium for instrumental work, subject to license and infrastructure requirements.
- Run blind comparisons. Generate several tracks per prompt and randomize model labels. Repeated samples help expose consistency problems that a single output can hide.
- Inspect narrow slices. Overall Elo averages performance across 17 genres. Listen to relevant samples because a model’s strengths in EDM, classical music or vocal production may differ sharply.
- Test detailed briefs. Include mood, structure, instrumentation, vocal delivery and mix characteristics so the evaluation reflects the controls a production system needs.
- Review operational constraints. Confirm API access, rate limits, data handling, content policies, licensing and output rights before integration.
The v1.1 launch data places Suno v6 first for both vocal and instrumental generation, identifies MiniMax Music 3.0 as the leading open-weights entry, and shows how little separates several other systems. For developers, the benchmark provides a sharper shortlist and a useful evaluation pattern: detailed prompts, blind pairwise review and separate tests for vocal and instrumental workloads.