LM Arena's Factuality Leaderboard Reveals OpenAI Leads While Meta Falls 13 Spots

Arena.ai launches a factuality-weighted leaderboard built on 2 million labeled LLM claims, reshuffling which models actually sit on top

·
·
LM Arena's Factuality Leaderboard Reveals OpenAI Leads While Meta Falls 13 Spots
AuthorArena.ai
Read2 min
  • New leaderboard: Arena.ai launches a factuality-weighted ranking combining human preference (75%) and automated fact-checking (25%) for Text and Search Arenas.
  • Massive dataset: Over 2 million LLM claims labeled from real conversations -- 1.3M+ from Text Arena and 700k+ from Search Arena.
  • OpenAI wins on factuality trends: Only provider consistently improving factuality over time; Anthropic leads on human preference but lags on factual accuracy.
  • Big movers: GPT-5.5 jumps 13 spots in Text Arena; Meta's Muse Spark drops 13 spots; GPT-5.5-search takes the top Search Arena factuality spot.
  • Methodology: Atomic claims are extracted, filtered for web-verifiability, and scored by search agents; abstaining from claims is never penalized, only confident falsehoods are.
  • Legal tasks worst, math best: Models are least factual on Legal and Government queries and most accurate on mathematical tasks, per Arena's industry breakdown.

Human preference has always been the currency of Arena.ai -- the platform formerly known as Chatbot Arena that now sits at the center of how the industry decides which model is best. But preference and truth are not the same thing, and Arena just made that gap impossible to ignore. The team has launched a factuality-weighted leaderboard, combining human votes with automated fact-checking at a scale that has never been attempted in a live evaluation platform.

The problem with pure preference

Arena's leaderboard has always worked by pitting two anonymous models head-to-head and letting a human pick the better response. A user enters one prompt, two anonymous models answer it, and the user votes for the better response. Arena then feeds those results into a Bradley-Terry rating system, similar to Elo for pairwise competitions. Over 6 million votes from users worldwide make it the largest crowdsourced LLM benchmark in 2026.

The catch? The system penalizes hallucinations, refusals, and verbose responses humans find annoying -- but it is less useful for measuring factuality, code correctness, or math. A model that sounds confident and comprehensive can win a human vote even if half its claims are wrong. Arena's new factuality layer is a direct attempt to fix that.

Two million claims, one new signal

To power these rankings, Arena labeled over 2 million claims made by LLMs in real-world conversations -- 1.3+ million from Text Arena and 700k+ from Search Arena, spanning roughly 130k Text Arena battles and 40k Search Arena battles. The methodology is worth understanding in detail, because it is more rigorous than most factuality benchmarks.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves