OpenAI's GPT-6 Astra Cracks Epoch's Hardest Math Benchmark in 14 Months

A math benchmark built to resist AI just fell. GPT-6 Astra cracked the final Tier 4 problem, closing out a 98 percent run in 14 months.

·
·
OpenAI's GPT-6 Astra Cracks Epoch's Hardest Math Benchmark in 14 Months
  • GPT-6 Astra solved the last unsolved FrontierMath Tier 4 problem, pushing the top score to 98%.
  • Tier 4 went from 5% to saturation in roughly 14 months since its July 2025 launch.
  • The final holdout, authored by Jay Pantone, was reportedly solved without the unintended shortcuts mathematicians flagged on earlier problems.
  • Runners-up trail significantly: GPT-5.6 Sol at 83.0% and GPT-5.6 Terra at 68.3%.
  • Caveat: OpenAI funded part of FrontierMath's development and has exclusive access to a portion of it.
  • Epoch is pivoting evaluation to FrontierMath Erdős, where Astra solved only 2 of 68 problems officially.

A benchmark that was supposed to hold up for years just fell. FrontierMath, the research-mathematics test suite from Epoch AI, has seen its hardest tier fully cleared, with OpenAI's newest frontier model solving the one problem that had resisted every prior system.

GPT-6 Astra scored 98% on FrontierMath Tier 4, cracking the only remaining unsolved problem, and Epoch now considers the benchmark saturated. The final holdout was authored by combinatorialist Jay Pantone, and unlike many earlier Tier 4 items, mathematicians did not report that the model had exploited an unintended shortcut to get there.

From 5% to saturation in fourteen months

When Epoch launched Tier 4 in mid-2025, the top model of the day cleared just 5% of it. Fourteen months later, the ceiling is essentially gone. The runners-up on the current leaderboard are not far behind either: GPT-6 Astra leads with 97.6%, followed by GPT-5.6 Sol at 83.0% and GPT-5.6 Terra at 68.3%.

Tier 4 was designed to be the wall. Epoch AI gathered more than seventy human mathematicians to write problems ranging from accessible Tier 1 questions to Tier 4 items that push into genuine research territory. That wall being behind us reshapes how the field talks about mathematical reasoning as a frontier capability.

What actually got solved

FrontierMath is not a word-problem set. The Tier 4 items are unpublished, research-grade questions, and evaluation runs programmatically rather than by grading proofs. Astra's run was part of a broader sweep where it also set records on Epoch's math, continual learning, and game-puzzles benchmarks. On the long-horizon coding benchmark MirrorCode, it ranks between Opus 4.7 and Fable 5.

OpenAI describes the model as a general step up rather than a math specialist. GPT-6 Astra pulls together years of work across pre-training, reinforcement learning, and alignment, and posts state-of-the-art results on computer use, browsing, software engineering, cybersecurity, science, and professional work. The math result is one data point in a broader capability jump.

The caveats worth knowing

Saturating Tier 4 is not the same as saturating mathematics, and a few asterisks matter before anyone declares the field solved:

  • Funding disclosure. Epoch AI, which runs FrontierMath, notes OpenAI funded its development and has exclusive access to part of it. That relationship has been contentious since the benchmark launched.
  • Harder tests already exist. A separate benchmark of unsolved Erdős problems saw Astra solve only 2 of 68 in its official run, rising to 5 with repeated attempts that cost over $220,000 in compute.
  • Compute intensity. Astra's headline numbers were run at maximum reasoning effort, which inflates both latency and token spend.
  • It is not dominant everywhere. On Humanity's Last Exam, Astra's 57.2% trails its own predecessor's 65.0%.

Where the frontier moves next

Epoch has already pivoted to harder ground. Their Open Problems track collects significant unsolved research questions with computationally verifiable solutions, and FrontierMath Erdős is a curated set of 68 Erdős problems formalized in Lean, where models must produce complete proofs or disproofs rather than just final answers. Astra took the top spot on the Erdős set too, but the low absolute score there is the more honest indicator of where research math sits today.

Why this matters for your work

For anyone building on top of frontier models, the practical read is that closed-form, single-answer math questions no longer offer a meaningful moat, even at the research level. The interesting evaluation surface is shifting to:

  1. Long-horizon proof construction in formal systems like Lean, where the model has to build and verify a full argument.
  2. Genuinely open problems where no ground-truth answer exists yet.
  3. Multi-hour agentic tasks that mix computation, tool use, and error recovery, which one-shot academic tests do not capture.

If you have been using FrontierMath scores as a proxy for reasoning quality in model selection, that signal has now flatlined at the top. The next round of differentiation will come from benchmarks still sitting in the single-digit-percent regime, and that gap is where the honest capability picture lives.

Comments

avatar