Vals AI's MysteryMechanism Benchmark Shows GPT-6 Astra Tops Science Test at 53%
A new Vals AI benchmark forces frontier models to rediscover hidden math laws through experiments, exposing where extra reasoning stops paying off.
- Vals AI launched MysteryMechanism, a benchmark for rediscovering hidden math laws through bounded experiments.
- GPT-6 Astra leads at 53% accuracy, Claude Opus 5.5 second at 50%, across 17 models.
- Agents get anonymous variables, two passive observations, and a 2d+1 experiment budget with no domain hints.
- Cranking reasoning effort to max can cost 20x more with zero accuracy gain past each model's sweet spot.
- Astra runs experiments sequentially to rule out hypotheses; Opus tends to commit to one formula early.
- Around 90% of Astra's successful traces never named the underlying domain, suggesting real inference, not recall.
AI Reasoning Models Hit 53% in a Black-Box Science Test
Vals AI has released MysteryMechanism, a benchmark that asks AI agents to run experiments, infer an unknown equation, and submit an executable formula. GPT-6 Astra leads with 53.15% accuracy, followed by Claude Fable 5.1 at 47.75%. For agent developers, the results show how experimental strategy and reasoning settings affect accuracy, cost, and latency under a strict tool-call budget.
Black-box science on a budget
Each task gives the agent anonymous input variables, their physical bounds, two initial observations, access to a persistent shell, and a budget of 2d+1 experiments for a mechanism with d inputs. The agent receives no domain labels or internet access. It must choose input values, observe the outputs, and infer the function that connects them.
The dataset contains 222 mechanisms drawn from biology, physics, chemistry, engineering, ecology, and abstract dynamics. Examples include blood viscosity as a function of hematocrit and shear rate, and stream-power incision in geomorphology. The benchmark removes those names and presents each mechanism as raw tuples such as (x1, x2) -> y.
Vals evaluates each submitted expression on private input probes designed to distinguish functional structures. A task receives binary credit when its normalized error falls below a threshold that accounts for noise. Functionally equivalent expressions earn credit regardless of their symbolic form or the name of the source mechanism.
More compute reaches a ceiling
Vals tested GPT-6 Astra and Claude Opus 5.5 at every available reasoning level. Higher settings allocate more computation, raising token use, cost, and latency. The two models produced different accuracy-versus-cost curves:
- GPT-6 Astra improves through
xhigh. Itsmaxsetting delivers the same accuracy while costing 47% more and taking 50% longer. - Claude Opus 5.5 plateaus around
high. Moving fromhightoxhighroughly doubles the reasoning work and triples the cost with almost no accuracy gain. - Opus at
maxfinally tests different formulas, producing a measurable score increase after the earlier plateau.
Additional computation helps when it changes the hypothesis under test. The traces show little return when a model spends its larger budget repeatedly adjusting one incorrect formula. Across effort settings, per-task cost can increase by as much as 20-fold, making the location of each model’s accuracy plateau operationally significant.
Two search strategies split the field
At its lowest effort setting, Opus submits all experiments in one command, fits parameters for a single candidate equation, and proceeds to its answer without testing alternatives. Astra runs experiments sequentially and selects each new input to eliminate remaining candidates. That adaptive approach extracts more information from the same experiment allowance.
Astra at low matches Opus at high in accuracy at a similar cost, although Astra-low costs 3.7 times as much per task as Opus-low. Both models complete their tool calls and use their experiment budgets. Opus-low loses accuracy because it commits to an incorrect functional form more often, not because its tools fail.
Astra leads a low-scoring field
GPT-6 Astra solves 118 of the 222 mechanisms, yielding 53.15% accuracy. The selected public results below also report how often a successful trace left the source domain unnamed.
| Model | Accuracy | Domain unnamed in successful traces |
|---|---|---|
| GPT-6 Astra | 53.15% | 89.8% |
| Claude Opus 5.5 | 50% | Not reported |
| Claude Fable 5.1 | 47.75% | 74.5% |
| Claude Opus 5 | 37.39% | 79.5% |
| Gemini 3.8 Flash | 36.49% | 66.7% |
| Muse Spark 1.3 Max | 36.04% | 92.5% |
| GPT-5.6 Sol | 33.33% | 89.2% |
| DeepSeek Flash 4.1 | 21.17% | 97.9% |
In a blinded post-hoc audit, 89.8% of Astra’s successful traces left the correct source domain unnamed. Similar patterns appear for several other models. The result supports the view that successful agents can recover functional relationships through experimentation and curve fitting without explicitly recognizing the textbook context. The audit cannot establish whether latent familiarity with common equation forms influenced those answers.
A narrow test with practical signals
MysteryMechanism isolates a specific part of scientific and engineering work: selecting informative inputs and recovering a compact mathematical function from sparse observations. That capability also appears in debugging, system identification, performance tuning, and data analysis, where an agent must choose the next test instead of receiving a complete problem statement.
The controlled setup omits literature review, instrument failure, ambiguous measurements, collaboration, and validation against external evidence. Its accuracy figures are task success rates under a fixed query budget and offer no general measure of scientific competence.
How to choose a reasoning tier
Agent builders should tune reasoning effort against representative workloads instead of assigning the highest setting by default. Opus 5.5 scores 69.69% on the broader Vals Index, yet its MysteryMechanism traces show that aggregate performance does not determine the best configuration for exploratory search.
- Measure success rate, cost, and latency at every available reasoning tier.
- Use the last tier that produces a meaningful accuracy gain as the default.
- Inspect traces for repeated parameter fitting without changes to the underlying hypothesis.
- Reserve expensive escalation for tasks whose value justifies the added cost.
For these benchmark tasks, Astra’s efficient stopping point is xhigh. Opus reaches its first plateau at high; reaching its later improvement requires paying for max, where the model begins exploring alternative formulas.
Vals publishes the full leaderboard and traces on the benchmark page. Additional behavior and evaluation results appear on the Astra model card and Opus model card.