Epoch Expands FrontierMath to 50 Unsolved Problems AI Has Already Cracked Three

Epoch AI expands FrontierMath: Open Problems to 50 genuinely unsolved math problems, with AI already cracking three of them including a superpermutation record set by GPT-5.6

·
·
Epoch Expands FrontierMath to 50 Unsolved Problems AI Has Already Cracked Three
  • 50 unsolved problems: Epoch AI expanded FrontierMath: Open Problems from 14 to 50 genuinely unsolved research math problems.
  • 3 solved by AI: AI has solved three problems so far, all in the "Moderately Interesting" tier; harder tiers remain untouched.
  • Superpermutation record broken: GPT-5.6-Sol beat the best-known upper bounds for superpermutations over 8, 9, and 10 symbols in an interactive session.
  • Notability tiers: Problems are rated from "Moderately Interesting" to "Breakthrough"; the hardest include a faster prime factorization algorithm and the Smooth 4D Poincaré Conjecture.
  • OpenAI is the only verifier buyer: Verifier access is purchasable by anyone, but only OpenAI has bought in; Epoch owns Open Problems independently.
  • Context: FrontierMath Tier 1-3 scores jumped from under 2% to over 52% in two years, making a benchmark of unsolved problems the only durable yardstick left.

FrontierMath: Open Problems has grown from a 14-problem pilot to a full collection of 50 problems spanning combinatorics, number theory, algebraic geometry, and topology. These are problems that professional mathematicians have tried and failed to solve. Any AI solution would constitute a real advance in human mathematical knowledge.

A benchmark that bites back

The distinction from the original FrontierMath benchmark (Tiers 1–4) is fundamental. In the main benchmark, problems have known solutions created by experts. In Open Problems, professional mathematicians have attempted the problems and come up empty. What makes the collection tractable at scale is a design constraint: even though no solution is known today, each problem comes with a bespoke computer program, called a verifier, that can check whether a proposed answer is correct. Genuinely open, yet automatically checkable.

Problems are contributed by professional mathematicians who draw from their own research and rate how meaningful a solution would be. The benchmark organizes problems into four notability tiers:

  • Moderately Interesting — solid but narrow results within a subfield
  • Solid Result — publishable advances of broader interest
  • Major Advance — significant results that would attract wide attention
  • Breakthrough — results that would reshape a field

The hardest tier includes problems like finding a faster prime factorization algorithm and adapting Apéry's proof of the irrationality of ζ(3) to other constants. Estimated solving times range from one to four weeks at the low end to three to ten years at the high end. The number of serious human attempts per problem ranges from two or three mathematicians to over fifty.

Three down, forty-seven to go

At launch, AI has already solved three problems. The most dramatic is the superpermutation result. A superpermutation over n symbols is a sequence containing every possible ordering of those symbols as a consecutive subsequence — the shortest string that visits every permutation. The minimum length is known exactly only for small n, and beating the best-known upper bounds for n = 8, 9, 10 was an open problem.

On July 26, 2026, Uku Raudvere posted a superpermutation over 8 symbols shorter than the then-current upper bound to GitHub and to the Superpermutators Google group. Inspired by that post, William Echols posted superpermutations over 9 and 10 that same day, developed in an interactive session using OpenAI's GPT-5.6-Sol at extended reasoning. GPT-5.6 Sol then sourced those permutations and passed the verifier during a pre-release test, earning the benchmark's first superpermutation solve. Mathematicians are still working out whether the approach generalizes to an improved bound for all n.

The second solve is more subtle. The genus 2 Jacobian torsion problem asked for a genus 2 curve over the rational numbers with a rational torsion point of order 31 or higher. A qualifying curve had already been found by humans in a related project but had not been recognized as a solution to this specific problem, so it counts as solved and has been removed from the active benchmark.

The third solve — a Ramsey-style hypergraph problem — has a longer history. Kevin Barreto and Liam Price first elicited a solution using GPT-5.4 Pro. Problem contributor Will Brian confirmed it, and a write-up is in progress. Epoch subsequently determined the problem did not meet the minimum notability bar and removed it, so the official solved count on the active benchmark stands at two problems in the "Moderately Interesting" tier, with the harder tiers still untouched.

Rules of the game

A few design choices are worth understanding before you follow or use the benchmark:

  • Human-first removal: If a human solves a problem before AI does, it gets pulled. This already happened once — a "Major Advance" tier problem was solved by humans before launch and removed.
  • Notability pruning: Two problems were removed because their verifiers could not detect correct solutions with sufficient fidelity.
  • Verifier access is paid: Any party can purchase access to the verifiers, with proceeds funding benchmark expansion. The main cost is mathematician compensation, since formulating and implementing each problem is labor-intensive. OpenAI is currently the only entity to have purchased access.
  • Independent ownership: OpenAI funded the original FrontierMath Tiers 1–4, but Open Problems is developed and owned solely by Epoch.

Why the timing matters

The broader FrontierMath story is one of rapid escalation. At launch, every frontier model scored below 2% on the benchmark. By April 2026, OpenAI's GPT-5.5 Pro solved 52.4% of Tier 1–3 problems and 39.6% of Tier 4 problems — a more than 25-fold improvement in under two years. Open Problems exists because models are saturating the solved-problem benchmarks, and the field needs a benchmark that cannot be saturated by definition.

In mid-2024, high school math was still a genuine challenge for AI systems. By the end of 2025, those systems were solving problems designed to be tractable only for top human experts. The Open Problems benchmark is built to detect and measure the next step: problems no human has solved before.

Scores alone leave one question open. As Epoch notes in its benchmark overview, even the most impressive AI contributions to mathematics have so far applied known techniques. A genuinely novel theoretical contribution — deriving new mathematical machinery from scratch — would be a different kind of milestone, and one the benchmark is designed to help identify through post-hoc analysis of solutions.

What to watch

The problem statements and prompts are public, and the data is available for download. Running a model against the verifiers requires purchasing access from Epoch; currently only OpenAI has done so. The editorial board — mathematicians Thomas Bloom, Little Math, and Dan Romik — vets each problem for notability and correctness.

The unsolved problems in the higher tiers include the Smooth Four-Dimensional Poincaré Conjecture, a faster prime factorization algorithm, and Apéry-style irrationality proofs. Solving any one of them would be front-page news in the math world, and AI is now being seriously evaluated against all of them.

Comments

avatar