Artificial Analysis Rebuilds Its Image Arena to Rank Models by Real Work

Artificial Analysis rebuilds its image arena around 10 real-world use cases and 9 model capabilities, revealing where GPT Image 2's dominance ends and cheaper models take over

·
·
Artificial Analysis Rebuilds Its Image Arena to Rank Models by Real Work
  • Artificial Analysis has overhauled its Text to Image Arena to rank models across 10 use cases and 9 capabilities, not just a single overall score.
  • GPT Image 2 (Elo 1339) leads every sub-leaderboard, but costs $211 per 1,000 images vs. $67 for Nano Banana 2.
  • Reve 2.1 (Elo 1299) holds #2 overall and ties GPT Image 2 on UI/UX; it is a 4K layout-first model from a ~65-person lab trained on 10x fewer GPUs.
  • Nano Banana 2 is the value pick for most use cases including UI/UX, Text Rendering, and Animation & Gaming at one-third the price of GPT Image 2.
  • Physics is the weakest capability across all top-10 models; no model ranks it as its strongest area.
  • Prompts are refreshed monthly from anonymized real-user data and retired when they stop discriminating between models, preventing leaderboard overfitting.

Most image generation benchmarks answer one question: which model is best? That question is increasingly useless. A product team prototyping UI mockups, a game studio generating concept art, and a marketing agency building campaign visuals all need different things from a model. Artificial Analysis has overhauled its Text to Image Arena to answer the more useful question: which model is best for your specific work?

A taxonomy, not a single score

The updated evaluation now measures models across a structured taxonomy of 10 real-world use cases and 9 model capabilities. Text-to-image models vary in capability across a wide range of use cases. One model may lead on text rendering, another on complex layouts or photorealistic human characters. The new leaderboard makes that variation explicit instead of burying it under a single aggregate score.

The use cases covered include:

  • Marketing & Advertising
  • Retail & E-commerce
  • Live-Action Film
  • Animation & Gaming
  • Architecture & Real Estate
  • Productivity & Knowledge Work
  • UI/UX Design
  • Social Media & Creator Content
  • Consumer
  • Frontier

The capability axes are equally granular:

  • Reasoning, Knowledge, Text Rendering, Layout
  • Complex Compositions, Lighting, Material
  • Physics, Human Anatomy

How the scoring works

The methodology measures model quality through human preference across this structured taxonomy. Prompts are sampled evenly across every use case and capability combination, so each contributes equal weight to the overall leaderboard rating.

The arena runs blind pairwise votes: models are ranked using an Elo rating system derived from user votes in blind comparisons. Users compare two images generated from the same prompt without knowing which model created each image. A few additional design choices make the signal cleaner:

  • Judging hints surface short cues tied to the prompt's use case and capability, directing voter attention to what actually matters on complex prompts

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves