Artificial Analysis Patches Intelligence Index to Crown Claude Opus 5
Artificial Analysis patches its Intelligence Index to v4.1.1, upgrading grader models and fixing scoring errors in the τ³-Banking benchmark

- Artificial Analysis patched its Intelligence Index to v4.1.1, fixing grading errors without changing benchmark weights or rankings.
- τ³-Banking updated to v1.0.1, correcting scoring bugs in agent trajectories that recover from errors on multi-step banking tasks.
- HLE, AA-LCR, and AA-Omniscience now all use GPT-5.6 Luna as grader, replacing three different legacy models for consistency.
- Claude Opus 5 remains #1 with an Intelligence Index score of 63; most models move less than 1 point.
- Muse Spark 1.2 (xhigh) saw the largest gain at +2.7 points due to improved grading robustness.
- All scores on artificialanalysis.ai now reflect v4.1.1 automatically; free to access.
Artificial Analysis has patched its Intelligence Index from v4.1 to v4.1.1. The scope is narrow by design: no new benchmarks, no reshuffled weights. The fixes target grading errors and grader inconsistency across evaluations. Claude Opus 5 holds the top spot with an Index score of 63.
What changed
Two categories of fixes landed in this patch. The τ³-Banking benchmark was updated to its v1.0.1 release from Sierra Platform, and three evaluations received new grader models.
The v1.0.1 grading update corrects errors in the banking_knowledge task, specifically in trajectories where an agent recovers from a mistake mid-task. These "unhappy paths" were previously scored incorrectly. Scores from earlier versions are not directly comparable to v1.0.1 results.