Artificial Analysis Rebuilds AA-Video to Rank Models by Real-World Jobs
Artificial Analysis rebuilt its text-to-video benchmark around 1080p playback, 1,000 tagged prompts, and 68,000 human votes across 10 use cases and 10 capabilities.
- Artificial Analysis launched AA-Video-T2V v2.0 and a silent variant, judging every model at 1080p.
- 68,000+ blind human votes on 1,000 prompts tagged by use case, capability, and visual style.
- Wan 3.0 ranked #1 at launch, leading 10 of 20 category boards at $12/min.
- MiniMax H3 puts an open-weights model on the Pareto frontier at $4.80/min for 768p.
- Gemini Omni Flash 1.1 dominates audio synchronization and UI/UX motion design categories.
- Full methodology covers -18 LUFS audio leveling, 10 Mbps H.264, and per-session quality checks.
AA-Video v2 ranks video models by the jobs they handle
Artificial Analysis has released AA-Video-T2V v2.0, a pair of human-preference benchmarks for generated video. The benchmarks add separate rankings for production use cases, technical capabilities, and visual styles, with standardized playback capped at 1080p. Developers can now compare models for tasks such as animation, camera control, text rendering, and lip-synced dialogue instead of relying on one aggregate score.
One benchmark becomes two
The release separates evaluation into AA-Video-T2V v2.0, which scores text-to-video generation with audio, and AA-Video-T2V-Silent v2.0, which scores silent output. The audio benchmark contains 1,000 prompts, while the silent benchmark contains 500.
Both prompt pools are refreshed regularly. Artificial Analysis retires prompts when they stop distinguishing among models or no longer reflect current prompting patterns.
Each prompt carries tags across three dimensions: a production use case, a technical capability, and a visual style. Sampling is balanced across use-case and capability combinations. A single comparison can therefore contribute to several category rankings as well as the overall Elo score.
What each score contains
| Dimension | Categories |
|---|---|
| Use case | Marketing & Advertising; Retail & E-commerce; Live-Action Film; Animation & Gaming; Architecture & Real Estate; Productivity & Knowledge Work; UI/UX & Motion Design; Consumer; Social Media & Creator Content; and Frontier tasks that stretch current model capabilities. |
| Capability | Physics; Complex Composition; Human Anatomy; Lighting & Materials; Multi-Scene & Narrative; Camera Control; Text Rendering; Spatio-Temporal Consistency; Audio Synchronization; and Dialogue & Lip Sync. |
| Style | Visual treatments such as Cartoon & Anime. |
The overall Elo summarizes performance across the full prompt mix, rewarding models that handle varied tasks consistently. Category boards reveal specialists whose performance may be obscured by an aggregate ranking.
Elo is a relative rating derived from head-to-head results. A model gains points when evaluators prefer its output over another model’s response to the same prompt. Ratings can move as votes accumulate, and models with similar scores may be statistically tied.
Playback rules remove easy advantages
Evaluators watch two clips generated from the same prompt without seeing the model names. Both clips play at full size, and voting remains disabled until each has played for at least eight seconds.
- Resolution: Outputs retain their generated resolution up to 1080p. Higher-resolution clips are downscaled to 1080p, while lower-resolution clips are never upscaled.
- Encoding: Videos use H.264 with bitrate peaks capped at 10 Mbps.
- Audio: Clips are adjusted toward -18 LUFS using one fixed gain, limiting any advantage from higher playback volume.
- Quality control: Roughly one in five matchups has a strong consensus answer. Artificial Analysis discards an entire evaluation session when it fails those checks.
Different boards produce different leaders
The launch analysis ranked Wan 3.0 first overall and credited it with leading 10 of 20 category boards. The results also showed substantial specialization among the highest-ranked models.
| Model | Launch result | Reported price |
|---|---|---|
| Wan 3.0 | First overall; led categories including Cartoon & Anime and Animation & Gaming | $12 per generated minute |
| Dreamina Seedance 2.5 | Second overall; led Human Anatomy and Dialogue & Lip Sync | $34.12 per generated minute |
| MiniMax H3, 768p | Statistically tied for third | $4.80 per generated minute |
| Gemini Omni Flash 1.1 | Fifth overall; led UI/UX & Motion Design and Audio Synchronization | Not stated |
A later live-board snapshot placed Gemini Omni Flash first in text-to-video with audio at 1233 Elo, followed by Wan 3.0 at 1229 and a Fal-tuned MiniMax H3 Max at 1227. MiniMax H3 was the highest-rated open-weights model at 1220. These values are snapshots rather than fixed results because new votes continually change the rankings.
The category gaps expose recurring failures
Text Rendering produced the widest score spread in the reported results. MiniMax H3 led the category, with a gap of more than 100 Elo points to Kling 3.0. Even highly rated outputs showed text degrading when the printed surface flipped or rotated.
Camera Control produced a tighter cluster because many models struggled with pans, dollies, orbits, and zooms. Wan 3.0 led that category by 45 Elo points, making it the clearest outlier among otherwise similar scores.
Native audio also changed model preferences. In comparisons that scored the same videos with and without sound, Gemini Omni Flash 1.1 gained 55 Elo points when audio was included, Seedance 2.5 gained 44, and FLUX 3 gained 38. Those gains are relevant to workflows involving ambience, sound effects, or synchronized speech.
Cost changes the model choice
The launch data reported a statistical tie among four frontier live-action models whose prices ranged from $4.80 to $34.12 per generated minute, a roughly 7.1-fold difference. Category scores and pricing can therefore produce a different shortlist from the overall ranking.
- Animation and gaming: Wan 3.0 led the corresponding use-case board.
- Human performance and dialogue: Seedance 2.5 led Human Anatomy and Dialogue & Lip Sync.
- UI motion and synchronized audio: Gemini Omni Flash 1.1 performed well above its aggregate score in those categories.
- Open-weights deployment: MiniMax H3 was the highest-rated open model in the later snapshot, though the cited variant generated at 768p.
- Camera-heavy sequences: Wan 3.0 held the largest reported lead in Camera Control.
Benchmark scores cover human preference under controlled playback. Production evaluations still need representative prompts and measurements for latency, failure rate, effective cost, resolution, API limits, reproducibility, licensing, moderation, and available generation controls.
Current scores and category boards are available on the live leaderboard. Artificial Analysis documents prompt handling, playback, voting, and quality controls in its full methodology.