Epoch Expands FrontierMath to 50 Unsolved Problems AI Has Already Cracked Three
Epoch AI expands FrontierMath: Open Problems to 50 genuinely unsolved math problems, with AI already cracking three of them including a superpermutation record set by GPT-5.6

- 50 unsolved problems: Epoch AI expanded FrontierMath: Open Problems from 14 to 50 genuinely unsolved research math problems.
- 3 solved by AI: AI has solved three problems so far, all in the "Moderately Interesting" tier; harder tiers remain untouched.
- Superpermutation record broken: GPT-5.6-Sol beat the best-known upper bounds for superpermutations over 8, 9, and 10 symbols in an interactive session.
- Notability tiers: Problems are rated from "Moderately Interesting" to "Breakthrough"; the hardest include a faster prime factorization algorithm and the Smooth 4D Poincaré Conjecture.
- OpenAI is the only verifier buyer: Verifier access is purchasable by anyone, but only OpenAI has bought in; Epoch owns Open Problems independently.
- Context: FrontierMath Tier 1-3 scores jumped from under 2% to over 52% in two years, making a benchmark of unsolved problems the only durable yardstick left.
FrontierMath: Open Problems just grew up. Epoch AI has expanded its benchmark of genuinely unsolved research mathematics from a 14-problem pilot to a full collection of 50 problems, spanning combinatorics, number theory, algebraic geometry, and topology. These are not exam questions with known answers. They are problems that professional mathematicians have tried and failed to solve, and any AI solution would constitute a real advance in human mathematical knowledge.
The benchmark that bites back
The distinction between this and the original FrontierMath benchmark (Tiers 1-4) is fundamental. Unlike the main benchmark, where problems have known solutions that an expert created, these are problems that professional mathematicians have attempted and failed to solve. Problems in FrontierMath: Open Problems are designed so that, even though no solution is known today, potential solutions can be checked for accuracy by a bespoke computer program, called a verifier. That combination -- genuinely open, yet automatically checkable -- is what makes the benchmark tractable at scale.
Problems are contributed by professional mathematicians who suggest problems from their own research and rate the meaningfulness of solutions, ranging from results of moderate interest within a subfield all the way up to major breakthroughs. The benchmark organizes problems into four notability tiers:
- Moderately Interesting -- solid but narrow results within a subfield
- Solid Result -- publishable advances of broader interest
- Major Advance -- significant results that would attract wide attention
- Breakthrough -- results that would reshape a field
The hardest tier includes problems like finding a faster prime factorization algorithm and adapting Apéry's famous proof of the irrationality of ζ(3) to other constants -- problems that, if solved, would be landmark results in number theory. Estimated solving times range from one to four weeks at the low end to three to ten years at the high end, and the number of serious human attempts per problem ranges from two or three mathematicians to over fifty.