LM Arena Merges 12 AI Leaderboards Into One Free Dashboard

Arena rolls out a redesigned leaderboard hub that unifies live model rankings, category standings, first impressions, and news in one view.

·
·
  • Arena launched a redesigned Leaderboard Overview unifying rankings, live sessions, and news.
  • New Release Rankings surface fresh models like Claude Opus 5.5 and GPT-6 Sol with category standings.
  • Performance by Category covers Agents, Coding, WebDev, Text, Image, and Video top-10s.
  • Pareto frontier view plots agent score against dollar cost per task for price-performance analysis.
  • Model Capabilities section adds short video first impressions from the Arena team.
  • Backed by 6.8M+ blind human votes across 360+ models, no login required.

Arena puts its model rankings on one page

Arena, the crowd-voted evaluation service formerly known as LMSYS Chatbot Arena, has launched a redesigned Leaderboard Overview. The dashboard combines recent releases, category leaders, live sessions, price-performance data, model walkthroughs, and Arena news, reducing the need to search across separate tabs.

Four modules in one view

The public dashboard requires no account or subscription. Arena reports more than 6.8 million blind human votes across over 360 models. Each vote records a preference between two model outputs, and the aggregate results produce relative rankings that change as more votes arrive. Positions for newly listed models are therefore provisional.

What the redesigned overview contains
Module What it shows
New Release Rankings A live feed of recently added models and their current category positions. At publication, the page listed Claude Opus 5.5 at No. 1 in WebDev and Recraft V4.1 Flash at No. 55 in Text-to-Image.
Performance by Category Top-10 tables and live sessions across Agents, Coding, WebDev, Text, Image, and Video. The Agents view also plots performance against cost per task.
Model Capabilities Short video walkthroughs featuring models such as GPT-6 Sol, GPT-6 Astra, Claude Fable 5.1, Qwen 3.8 27B, and Kimi K3. These videos contain the Arena team’s early observations.
Arena News Posts covering methodology changes, evaluation frameworks, and academic partnerships.

Twelve boards, one entry point

  • Text, Agent, and WebDev
  • Image-to-WebDev, Text-to-Image, and Image Edit
  • Text-to-Video, Image-to-Video, and Video Edit
  • Vision, Document, and Search

The former single-board layout obscured specialized results as Arena added categories. The overview now exposes those boards from one page, making it easier to compare models within the workload that matters.

Arena Expert has also added another layer to the rankings. The framework selects 5.5% of Arena prompts based on reasoning depth and specificity, creating a smaller evaluation set designed to separate models that score similarly on the broader leaderboard.

Cost joins the comparison

The Agents module plots performance against dollar cost per task and draws a Pareto frontier. A model lies on that frontier when no displayed alternative is both cheaper and higher-scoring, giving teams a direct view of the trade-off between capability and operating cost.

Price-performance examples shown at publication
Model Displayed performance Cost per task
Tencent Hy3 5.23% $0.04
Claude Fable 5.1 (Max) 13.71% $4.15

The figures are snapshots and may change as Arena updates its results. Their placement still gives developers a quick way to identify lower-cost candidates before running workload-specific tests.

Using the dashboard in practice

  • Release triage: Find recently added models and check their initial category positions.
  • Task selection: Compare leaders in Coding, WebDev, Agents, Image, Video, and other specialized boards.
  • Budget planning: Examine agent performance alongside estimated cost per task.
  • Qualitative review: Inspect live sessions and walkthroughs before conducting controlled tests.

Arena AI provides no public API, so teams that need stable, machine-readable data must use another source. Community projects such as arena-ai-leaderboards provide historical snapshots, although unofficial mirrors may lag site changes. Research workflows should record retrieval dates and pin the exact dataset used.

Close ranks need context

Independent trackers report that Arena’s top 10 models sit within roughly 20 Elo points of one another. Such a narrow spread leaves the ordering sensitive to additional votes and changes in the prompt mix. The scores’ predictive value also depends on how closely Arena’s voters and prompts resemble a team’s production workload.

Category rankings provide a stronger starting point than the overall table for most engineering decisions. Final selection still requires controlled tests using representative prompts, latency targets, output constraints, and expected request volumes.

Trending
  • No trending articles

Comments

avatar

Next Reads