LM Arena Breaks Down 52 Image Models Across 7 Categories With 28M Votes
Arena.ai adds seven granular categories to its image editing leaderboards, revealing GPT Image 2 leads every dimension across 28M+ human votes

- Arena.ai launched seven granular categories (Photorealistic, Portraits, Cartoon, Art, 3D, Commercial Design, Text Rendering) for its Single and Multi-Image Edit leaderboards.
- GPT Image 2 ranks #1 in all seven categories, with its largest margin in Text Rendering; the leaderboard covers 52 models and 28M+ votes.
- For Single-Image Edit, Meta's Muse Image (#2) and Microsoft's MAI-Image-2.5 (#3) are closely matched behind GPT Image 2.
- For Multi-Image Edit, Bytedance's Seedream 5.0 Pro and Meta's Muse Image are nearly tied, with divergence in 3D Modeling and Cartoon.
- Analysis of 3M+ votes found Photorealistic and Portraits overlap 64%, but Photorealistic and Art overlap under 1%, confirming they measure genuinely different capabilities.
- The full leaderboard is free at arena.ai/leaderboard/image-edit; historical data is on Hugging Face.
A single overall score has always been a blunt instrument for picking an image model. The best portrait editor might be mediocre at rendering logos, and the best 3D renderer might struggle with photorealism. Arena.ai's image editing leaderboard just got a lot more useful: the platform has rolled out seven granular categories for both its Single-Image Edit and Multi-Image Edit arenas, letting you see exactly where each model wins and loses.
What changed
Until now, Arena's image editing leaderboards showed one aggregate Elo score per model. The new update mirrors the category structure already live in the Text-to-Image arena, bringing the same seven buckets to editing:
- Photorealistic & Cinematic Imagery
- Portraits
- Cartoon, Anime & Fantasy
- Art
- 3D Imaging & Modeling
- Product, Branding & Commercial Design
- Text Rendering
The leaderboard now covers 52 models and has accumulated over 28 million votes. Each category gets its own ranked list, so you can filter by the exact output type you care about rather than trusting a single averaged number.
How the scoring works
Arena uses a blind pairwise voting system. Users submit a prompt, receive outputs from two anonymous models, and choose the better one. The platform aggregates millions of such votes using a Bradley-Terry-Luce model to estimate per-model strength scores, reported as Elo-style numbers. Think of it like chess ratings: a 10-point Elo gap means one model wins roughly 55-60% of head-to-head matchups.