LM Arena's Factuality Leaderboard Reveals OpenAI Leads While Meta Falls 13 Spots

Arena.ai launches a factuality-weighted leaderboard built on 2 million labeled LLM claims, reshuffling which models actually sit on top

·
·
LM Arena's Factuality Leaderboard Reveals OpenAI Leads While Meta Falls 13 Spots
  • New leaderboard: Arena.ai launches a factuality-weighted ranking combining human preference (75%) and automated fact-checking (25%) for Text and Search Arenas.
  • Massive dataset: Over 2 million LLM claims labeled from real conversations -- 1.3M+ from Text Arena and 700k+ from Search Arena.
  • OpenAI wins on factuality trends: Only provider consistently improving factuality over time; Anthropic leads on human preference but lags on factual accuracy.
  • Big movers: GPT-5.5 jumps 13 spots in Text Arena; Meta's Muse Spark drops 13 spots; GPT-5.5-search takes the top Search Arena factuality spot.
  • Methodology: Atomic claims are extracted, filtered for web-verifiability, and scored by search agents; abstaining from claims is never penalized, only confident falsehoods are.
  • Legal tasks worst, math best: Models are least factual on Legal and Government queries and most accurate on mathematical tasks, per Arena's industry breakdown.

Human preference has always been the currency of Arena.ai -- the platform formerly known as Chatbot Arena that now sits at the center of how the industry decides which model is best. But preference and truth are not the same thing, and Arena just made that gap impossible to ignore. The team has launched a factuality-weighted leaderboard, combining human votes with automated fact-checking at a scale that has never been attempted in a live evaluation platform.

The problem with pure preference

Arena's leaderboard has always worked by pitting two anonymous models head-to-head and letting a human pick the better response. A user enters one prompt, two anonymous models answer it, and the user votes for the better response. Arena then feeds those results into a Bradley-Terry rating system, similar to Elo for pairwise competitions. Over 6 million votes from users worldwide make it the largest crowdsourced LLM benchmark in 2026.

The catch? The system penalizes hallucinations, refusals, and verbose responses humans find annoying -- but it is less useful for measuring factuality, code correctness, or math. A model that sounds confident and comprehensive can win a human vote even if half its claims are wrong. Arena's new factuality layer is a direct attempt to fix that.

Two million claims, one new signal

To power these rankings, Arena labeled over 2 million claims made by LLMs in real-world conversations -- 1.3+ million from Text Arena and 700k+ from Search Arena, spanning roughly 130k Text Arena battles and 40k Search Arena battles. The methodology is worth understanding in detail, because it is more rigorous than most factuality benchmarks.

The pipeline works like this:

  1. A random sample of battles is selected for a factuality audit.
  2. Each model response is broken into atomic claims -- individual, independently verifiable statements.
  3. Claims that are subjective, ambiguous, or not web-verifiable are filtered out (opinions, fictional context, etc.).
  4. A system of search agents verifies each remaining claim and assigns a calibrated truth probability.
  5. The model with higher average claim truthfulness wins the factuality battle.

In Text Arena battles, a claim was found in at least one response 76% of the time; in Search Arena 88% of the time. In Text Arena, models averaged 5 claims per response; in Search Arena models averaged nearly 10. The marginal true claim rate in Text Arena was 87%, and 89% in Search Arena -- assisted by Arena's AutoModality classifier, which automatically routes prompts that likely need web access to Search Arena.

The math behind the composite score

Arena uses what it calls a composite Bradley-Terry model -- a way of fitting a single rating vector that simultaneously explains both human votes and factuality outcomes. The default factuality weight is set to 25%, meaning three-quarters of a model's score still comes from human preference. The weight is adjustable via a slider in the leaderboard UI, letting you dial from pure preference (0%) to pure factuality (100%).

One subtle but important design choice: abstaining from making a claim is never penalized. Only confident false claims count against a model. This prevents the system from rewarding models that dodge every question to stay technically accurate.

Who moves, and who falls

The ranking shifts when factuality is switched on are significant. In the Text Arena:

  • GPT-5.5 saw the largest gain, jumping up 13 spots into the #7 position.
  • Claude Fable 5 slips slightly to #2.
  • Muse Spark (Meta) dropped the hardest, falling 13 spots from #7 to #20.
  • By lab, Meta saw the largest overall drop (#2 to #5), while Anthropic held the #1 spot.
  • Among open model providers, Xiaomi saw the biggest improvement, jumping from #9 to #6.

In the Search Arena:

  • GPT-5.5-search moved up to take the top factuality spot.
  • GPT-5.2-search jumped from #11 to #3.
  • claude-sonnet-4-6-search fell from #6 to #9.
  • gemini-3.1-pro-grounding dropped from #7 to #13.

The bigger pattern: OpenAI is the only lab consistently improving factuality over time

Perhaps the most striking finding is the longitudinal trend. Arena plotted each provider's pure factuality score against flagship model release dates, and the picture is unambiguous.

While human preference scores across nearly all providers have been climbing over time, plotting pure factuality score against flagship model release date shows OpenAI is the only provider also consistently improving factuality over an extended period. SpaceXAI and Meta have also been improving with their recent releases.

The blog post adds important nuance by provider:

  • Anthropic: Models are generally more preferred by humans, but OpenAI's newest models are generally more factual. The most factual Claude before the newest 4.8 release was Claude 4.5.
  • Google: Gemini-2.5 models remain the most factual in the Gemini family, but Google appears to be getting less factual over time -- a notable regression.
  • SpaceXAI (Grok): Grok-4-0709 was solidly factual, but the subsequent Grok-4.1 series heavily regressed. Newer versions (4.20, 4.3, 4.5) have recovered.
  • Open source: Most open source models decline in score under higher factuality weight, most notably nvidia-nemotron-3-ultra. Exceptions are mistral-medium-3.5 and Tencent's huyuan-hy3-preview.

Why this matters beyond the rankings

The factuality and human preference signals turn out to be largely independent of each other. Arena found only a weak positive correlation between the two -- which is exactly why adding factuality as a separate axis is valuable. A model can be maximally factual by refusing to answer anything, and a model can win human preference votes by being confidently wrong in an articulate way. Neither extreme is what you actually want.

Hallucination rate is now a core safety metric in 2026, not just a quality metric. A model that confabulates confidently in a medical or legal context creates real harm. Arena's data backs this up: models are least factual on Legal and Government tasks, while performing best on mathematical tasks, with Software, Medical, and Scientific tasks falling in between.

The factuality toggle is live now in both Text and Search Arenas as a non-default option. It is not replacing the existing preference-based leaderboard -- it is layered on top of it, giving teams a second lens that is harder to game. Arena raised a $1.7 billion valuation after its Series A round, and moves like this signal it is building toward something more than a community vote counter -- a credible, multi-dimensional audit layer for the entire frontier model ecosystem.

Trending
  • No trending articles

Comments

avatar

Next Reads