Artificial Analysis Reshuffles Its Intelligence Index to Separate Frontier Models

Artificial Analysis rolls out Intelligence Index v4.2 early, swapping GPQA Diamond for long-horizon agentic work and multi-thousand-page document reasoning.

·
·
Artificial Analysis Reshuffles Its Intelligence Index to Separate Frontier Models
  • Artificial Analysis released Intelligence Index v4.2, pulling forward parts of the v5 roadmap.
  • AA-Briefcase added at 15% weight: private multi-week agentic knowledge work with thousands of source files.
  • GDP.pdf added at 10% weight: reasoning across 4,592 PDF pages in 100 professional tasks.
  • GPQA Diamond removed from the composite for being saturated and multiple-choice.
  • AA-LCR upgraded to v1.1, SciCode regraded at v1.0.1, category weights rebalanced across ten evaluations.
  • Claude Fable 5.1 tops the new index at 57, GPT-6 Astra second at 55.

The leaderboard most developers cite when comparing frontier models just reshuffled. Artificial Analysis Intelligence Index v4.2 pulls forward parts of the planned v5 rollout, retires a saturated benchmark, and shifts weight toward private, long-horizon tasks that resemble actual professional work.

Artificial Analysis says the accelerated timeline exists because too many models were hitting the ceiling on the old suite. A leaderboard that can't separate top contenders stops being useful to anyone choosing a model for production.

What changed inside the index

The v4.2 composite remains a weighted average across four categories: Agents (30%), Coding (20%), Scientific Reasoning (20%), and General (30%). The ten constituent evaluations are AA-Briefcase, GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, AA-LCR v1.1, AA-Omniscience, Humanity's Last Exam, GDP.pdf, and CritPt.

  • Added: AA-Briefcase (15% weight, Agents) and GDP.pdf (10% weight, General)
  • Removed: GPQA Diamond, dropped for saturation and because its multiple-choice format doesn't reflect real scientific work
  • Upgraded: AA-LCR bumped to v1.1 with a new system prompt, corrected answer keys, and GPT-5.6 Luna as the grader; SciCode regraded at v1.0.1
  • Rebalanced: weights redistributed across the ten evaluations to accommodate the two new additions

AA-Briefcase: agents on multi-week knowledge work

The headline addition is AA-Briefcase, 91 tasks spread across four scenarios, each modeled on multi-week knowledge work projects designed by industry experts. Agents run inside an offline E2B sandbox with a single shell and code-execution tool, get up to 500 turns per task, and receive realistic professional artifacts: Slack exports, spreadsheets, PDFs, interview transcripts, market research, board materials, and emails. They are then asked to produce deliverable output files.

Grading blends binary rubric checks with pairwise Elo comparisons on two axes: analytical quality (depth and structure) and presentation (professional readability). To reduce judge bias from within-family model preferences, rubric verdicts and pairwise comparisons rotate across three frontier judges: Claude Opus 4.8 at max effort, GPT-5.5 at high reasoning, and Gemini 3.1 Pro Preview at high reasoning. A final AA-Briefcase Elo aggregates all three signals via maximum-likelihood Elo fitting. The test set is private, which makes training-set contamination substantially harder to engineer.

GDP.pdf: reasoning across 4,592 pages of real documents

The second addition is Artificial Analysis' implementation of Surge AI's GDP.pdf. It covers 100 PDFs across ten professional domains, requiring models to synthesize evidence scattered across 4,592 pages of text, tables, charts, footnotes, and exclusions. Each task runs five times, producing 500 attempts per model graded against 1,275 criteria. The primary metric is All-pass: the share of the 500 attempts where every criterion passes, with GPT-5.6 Luna medium as the criterion judge.

Rather than routing raw PDFs through each provider's document-input API, Artificial Analysis extracts every page using LiteParse (with OCR where needed) and supplies ordered page images to vision-capable models. Provider document-input pipelines vary between runs and are largely opaque, which makes direct model comparisons unreliable without a standardized extraction layer.

Why GPQA Diamond was cut

Frontier models have essentially solved GPQA Diamond. The benchmark's four-option multiple-choice format also doesn't reflect how scientists actually work, so it offered two reasons to drop it simultaneously. Artificial Analysis will still run it on new releases and report a standalone number, but it no longer feeds the composite score.

How the leaderboard shifts

Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback) currently leads with an Intelligence Index score of 57, followed by GPT-6 Astra at 55. On GDP.pdf specifically, OpenAI leads with GPT-6 Astra at 33.2% and GPT-5.6 Sol at 28.2%, ahead of Claude Fable 5.1 at 26.2%. Those absolute numbers sit well below the near-ceiling scores common on saturated benchmarks, which is exactly what the redesign is trying to create: enough headroom for real model differences to register.

Three things to factor in if you use this index

  1. Scores don't carry over across versions. Models benchmarked under v4.1 versus v4.2 will look different not because the models changed but because the evaluation did. Any cross-version comparison needs to be rerun.
  2. Agentic capability now dominates. AA-Briefcase alone accounts for 15% of the composite, and the Agents category totals 30%. For single-turn Q&A workloads, the index maps less cleanly to real-world performance than it used to. Component scores matter more now.
  3. Long-context document reasoning is a first-class signal. Between GDP.pdf (100k+ token PDFs) and AA-LCR v1.1 (roughly 100k tokens per question across around 230 documents), the index now measures whether a model can reason across a stack of enterprise documents, not just retrieve chunks from them.

Benchmarking has a structural problem: once a test becomes a target, its signal decays. By expanding private test sets, anchoring agentic tasks to human-expert Elo baselines, and rotating multi-lab judge panels, Artificial Analysis is buying another cycle of useful differentiation before the frontier closes the gap again. More incremental updates are already planned ahead of the full v5 release.

Comments

avatar