Artificial Analysis Rebuilds Intelligence Index v5 With Private Coding Tests

Artificial Analysis is rolling out Intelligence Index v5 later this month, adding Terminal-Bench Science and a private-dataset coding benchmark to its composite leaderboard.

·
·
·
Artificial Analysis Rebuilds Intelligence Index v5 With Private Coding Tests
  • Artificial Analysis Intelligence Index v5 is scheduled to launch in late October.
  • v5 adds Terminal-Bench Science, an agentic eval of scientific workflows run through a terminal.
  • A new coding benchmark with a private test set joins the Coding category to reduce contamination risk.
  • Current Terminal-Bench Science leaders: GPT-6 Astra Max at 63.3%, Claude Opus 5.5 at 61.9%.
  • Category weights in v4.3.2 are Agents 30%, Coding 20%, General 30%, Scientific Reasoning 20%.
  • Existing Intelligence Index scores will not be directly comparable once v5 ships.

Independent model evaluator Artificial Analysis plans to release Intelligence Index v5 in late October, its largest revision of the composite benchmark this year. The update will add Terminal-Bench Science and a coding benchmark whose private test set will remain unpublished.

The Intelligence Index condenses results from several evaluations into a single model-capability score. Researchers already cite it in third-party model papers, so changing its benchmarks can reorder model rankings and break direct comparisons with earlier results.

V5 adds two tougher signals

Artificial Analysis has introduced the revision in stages. The current v4.3.2 build includes AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1.

Category Current weight
Agents 30%
Coding 20%
General 30%
Scientific reasoning 20%

The v5 release will add the following components:

  • Terminal-Bench Science: An existing standalone evaluation built by Stanford University researchers with the Terminal-Bench and Harbor team, with contributions from scientists at institutions worldwide. Its tasks reproduce research workflows performed through a command-line environment.
  • A private coding benchmark: An evaluation based on unpublished test data. Artificial Analysis has not disclosed its tasks, scoring method, or dataset composition.

Artificial Analysis has also yet to publish the final v5 category weights, an exact release date, or guidance for comparing v5 scores with v4.3.2 results. Those details are expected with the changelog.

Science work moves into the index

Terminal-Bench Science measures whether an AI agent can use a shell to complete realistic scientific workflows. The benchmark contains 70 tasks across five domains, and each task is attempted three times. Results are reported as pass@1: an attempt succeeds only when every test passes, while verifier timeouts count as failures.

The current standalone leaderboard indicates how demanding the evaluation is:

  1. GPT-6 Astra (Max): 63.3%
  2. Claude Opus 5.5 (Xhigh, Default Fallback): 61.9%
  3. Claude Opus 5.5 (Max, Default Fallback): 59.0%
  4. Claude Sonnet 5.5: 53.0%

No listed system clears 64%, leaving substantial room between current agents and complete task reliability. Adding the benchmark could shift composite rankings toward models that sustain long-running terminal sessions, recover from execution errors, and complete multi-step research procedures.

Private tests reduce training leakage

Public benchmark questions can enter model-training corpora after circulating online. A model may then reproduce familiar solutions rather than demonstrate that it can solve unseen problems, weakening the benchmark as a measure of coding ability.

Artificial Analysis has increasingly used private test variants and filtering procedures to reduce that risk, including its move to AutomationBench-AA. The new coding benchmark extends the same approach. Its unpublished data may provide a cleaner signal, although outside teams will be unable to reproduce the full evaluation independently or inspect the test distribution.

Prepare for a score break

  1. Rerun internal comparisons after release. V4.3.2 and v5 scores should be treated as separate series until Artificial Analysis publishes the new weights and methodology.
  2. Update coding-agent baselines. Terminal-Bench 4.0 already replaces v2.1 with a harder 66-task set, recalibrated compute and time allowances, and revised environments and verifiers. Scores may fall when teams move from the older harness.
  3. Track task-level scientific failures. Aggregate scores can hide whether an agent failed through planning, shell use, dependency management, computation, or verification. Those distinctions matter when improving research agents.
  4. Record benchmark versions. Reports should identify the Intelligence Index version, model configuration, and evaluation date so later results remain interpretable.

Artificial Analysis documents its current scoring and evaluation practices on its benchmark methodology page. The v5 changelog should clarify the final weights, private coding benchmark design, and rules for comparing results across index versions.

Trending
  • No trending articles

Comments

avatar

Next Reads