Anthropic's Claude Opus 4.8 Breaks Records on the Hardest AI Benchmark

Claude Opus 4.8 hits 1.5% on ARC-AGI-3, a benchmark where humans score 100% and every prior frontier model scored below 1%

·
·
Anthropic's Claude Opus 4.8 Breaks Records on the Hardest AI Benchmark
  • New SOTA: Claude Opus 4.8 scores 1.5% on ARC-AGI-3, up from the previous best of 0.43% by GPT-5.5.
  • Humans still dominate: Humans score 100% on ARC-AGI-3; the gap between AI and human performance remains enormous.
  • Better abstraction: Opus 4.8 perceives game environments as objects and systems, not pictures — a qualitative leap over Opus 4.7.
  • New failure mode exposed: Opus 4.8 clears early levels then commits to wrong sub-goals, burning hundreds of actions on contradictory theories.
  • ARC-AGI-2 scores strong: Opus 4.8 High reaches 72.08% on ARC-AGI-2 at $2.74/task; ARC-AGI-1 High hits 92%.
  • Available now: Opus 4.8 is live on claude.ai and the Claude API at $5/M input, $25/M output — same price as Opus 4.7.

Anthropic's Claude Opus 4.8 just became the new state-of-the-art on ARC-AGI-3, the hardest active benchmark in AI research, scoring 1.5% on the semi-private evaluation set. That number sounds tiny, but it's the highest any frontier language model has ever achieved on a test where humans score 100% and the previous best LLM sat at 0.43%.

What ARC-AGI-3 actually is

ARC-AGI-3 is an interactive reasoning benchmark that challenges AI agents to explore novel environments, acquire goals on the fly, build adaptable world models, and learn continuously. That's a meaningful departure from every benchmark that came before it.

Instead of presenting static puzzles with clear input-output pairs, it drops AI agents into interactive environments with no instructions, no stated goals, and no explicit rules. The agent has to figure out everything on its own through trial and observation, the same way a person would when handed a game they have never seen before.

ARC-AGI-3 represents hundreds of original turn-based environments, each handcrafted by a team of human game designers. There are no instructions, no rules, and no stated goals. To succeed, an AI agent must explore each environment on its own, figure out how it works, discover what winning looks like, and carry what it learns forward across increasingly difficult levels.

Why does this matter for benchmarking? ARC-AGI-3's interactive format is designed to resist the benchmark saturation cycle in a fundamental way. Because the environments are interactive rather than static, they cannot be memorized. Because there are no stated rules or goals, there is no fixed prompt structure to optimize against. Because the scoring system penalizes inefficiency so heavily, brute-force approaches that work by trying many strategies and keeping the one that works are essentially useless.

The gap is enormous, and that's the point

ARC-AGI-3 environments are difficulty-calibrated via extensive testing with human test-takers. Testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of the benchmark's launch, scored below 1%.

For context, these same frontier models dominate most other AI benchmarks. ARC-AGI-1 is approaching saturation, with top models now scoring well above 90%. The drop to sub-1% on ARC-AGI-3 suggests that the capabilities being tested here are different from what current models do well.

To put the Opus 4.8 result in historical context, here's where frontier LLMs stand on ARC-AGI-3:

  • Opus 4.8 (new SOTA): 1.5% , ~$10K cost
  • GPT-5.5: 0.43%
  • Opus 4.7: 0.18%
  • Most other frontier models at launch: below 0.51%
  • Humans: 100%

What Opus 4.8 does differently

The ARC Prize team didn't just record a score , they replayed every action and reasoning trace. What they found is qualitatively interesting. The biggest shift from Opus 4.7 is one of abstraction level: Opus 4.8 perceived environments as objects and systems, not as pixel grids or pictures. That's a higher-order representation, and it showed up in measurable ways.

On the game ar25, Opus 4.8 derived the Level 1 reflection rule by frame 5

Trending
  • No trending articles

Comments

avatar

Next Reads