Artificial Analysis Patches Intelligence Index to Crown Claude Opus 5

Artificial Analysis patches its Intelligence Index to v4.1.1, upgrading grader models and fixing scoring errors in the τ³-Banking benchmark

·
·
Artificial Analysis Patches Intelligence Index to Crown Claude Opus 5
Read1 min
SubtopicSmall Models
  • Artificial Analysis patched its Intelligence Index to v4.1.1, fixing grading errors without changing benchmark weights or rankings.
  • τ³-Banking updated to v1.0.1, correcting scoring bugs in agent trajectories that recover from errors on multi-step banking tasks.
  • HLE, AA-LCR, and AA-Omniscience now all use GPT-5.6 Luna as grader, replacing three different legacy models for consistency.
  • Claude Opus 5 remains #1 with an Intelligence Index score of 63; most models move less than 1 point.
  • Muse Spark 1.2 (xhigh) saw the largest gain at +2.7 points due to improved grading robustness.
  • All scores on artificialanalysis.ai now reflect v4.1.1 automatically; free to access.

Artificial Analysis has released a patch to its Intelligence Index, bumping it from v4.1 to v4.1.1. The update is deliberately narrow: no new benchmarks, no reshuffled weights. The goal is to keep the existing evaluation suite as reliable as possible by fixing grading errors and unifying the models that judge model outputs. Claude Opus 5 holds the top spot with an Index score of 63.

What actually changed

Two categories of fixes landed in this patch. First, the τ³-Banking benchmark was updated to its v1.0.1 release from Sierra Platform. Second, three evaluations got new grader models.

On the benchmark side, the v1.0.1 grading update fixes a couple of banking_knowledge task errors, and scores on that domain change as a result -- results produced with earlier versions are not directly comparable. Concretely, the fix targets trajectories where an agent recovers from an error mid-task (an "unhappy path"), which were previously being graded incorrectly.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves