EurekaBench Reveals AI Agents Still Fall 27 Points Behind Humans on Real Science
A new benchmark shows frontier agents can match human predictive accuracy on scientific problems but still fail to produce real understanding.
- EurekaBench evaluates AI agents on 26 expert-crafted scientific problems with 306 insight questions across six domains.
- Top agents (Claude Fable 5.1, GPT 6 Astra) hit 47.4% predictive accuracy, close to humans' 48.8%.
- But insight scores lag badly: 29.4-42.4% for agents vs 69.7% for human scientists.
- Most agents fail basic scientific-constraint tests on more than half the problems.
- Four recurring failures: unjustified choices, optimization-as-discovery, unresolved uncertainty, no follow-up on observations.
- Code and tasks available at github.com/EurekaBench/EurekaBench.
EurekaBench tests whether AI agents can explain the data they fit
Frontier AI agents came within 1.4 percentage points of the human baseline on EurekaBench’s predictive-accuracy metric. Their best scientific-insight score remained 27.3 points behind, and several systems violated basic domain constraints on more than half the tasks. The results suggest that benchmarks centered on prediction can overstate an agent’s ability to conduct scientific research.
Designed with ten doctoral-level scientists, the EurekaBench paper presents 26 research problems across neuroscience, geophysics, astrophysics, computer science, plasma physics, and chemistry. Each task supplies observations that challenge existing explanations, along with a simulator for testing hypotheses. The agent must submit a proposed mechanism, working Python code, and a record of its experiments.
Why a close fit can mislead
Most AI-scientist benchmarks emphasize accuracy on held-out data. A model can achieve that accuracy with correlations, tuned parameters, or flexible equations that reveal little about the system’s behavior. EurekaBench defines a mechanism as an executable account that respects domain knowledge, predicts unseen observations, and supports testable conclusions.
The benchmark scores submissions along three axes:
- Scientific constraints (SC): Does the mechanism respect expert-specified physical, biological, or computational assumptions?
- Predictive accuracy (PA): How accurately does the submitted Python implementation predict held-out data?
- Scientific insights (SI): Can a separate judge use the unchanged mechanism and additional simulator experiments to answer expert-authored questions about the system?
EurekaBench turns scientific insight into a measurable score through 306 yes-or-no questions. Domain experts wrote each question around a conclusion that a strong mechanism should support, such as how a system responds under a new condition or which process drives an observed effect.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.