Epoch Expands FrontierMath to 50 Unsolved Problems AI Has Already Cracked Three
Epoch AI expands FrontierMath: Open Problems to 50 genuinely unsolved math problems, with AI already cracking three of them including a superpermutation record set by GPT-5.6

- 50 unsolved problems: Epoch AI expanded FrontierMath: Open Problems from 14 to 50 genuinely unsolved research math problems.
- 3 solved by AI: AI has solved three problems so far, all in the "Moderately Interesting" tier; harder tiers remain untouched.
- Superpermutation record broken: GPT-5.6-Sol beat the best-known upper bounds for superpermutations over 8, 9, and 10 symbols in an interactive session.
- Notability tiers: Problems are rated from "Moderately Interesting" to "Breakthrough"; the hardest include a faster prime factorization algorithm and the Smooth 4D Poincaré Conjecture.
- OpenAI is the only verifier buyer: Verifier access is purchasable by anyone, but only OpenAI has bought in; Epoch owns Open Problems independently.
- Context: FrontierMath Tier 1-3 scores jumped from under 2% to over 52% in two years, making a benchmark of unsolved problems the only durable yardstick left.
FrontierMath: Open Problems has grown from a 14-problem pilot to a full collection of 50 problems spanning combinatorics, number theory, algebraic geometry, and topology. These are problems that professional mathematicians have tried and failed to solve. Any AI solution would constitute a real advance in human mathematical knowledge.
A benchmark that bites back
The distinction from the original FrontierMath benchmark (Tiers 1–4) is fundamental. In the main benchmark, problems have known solutions created by experts. In Open Problems, professional mathematicians have attempted the problems and come up empty. What makes the collection tractable at scale is a design constraint: even though no solution is known today, each problem comes with a bespoke computer program, called a verifier, that can check whether a proposed answer is correct. Genuinely open, yet automatically checkable.
Problems are contributed by professional mathematicians who draw from their own research and rate how meaningful a solution would be. The benchmark organizes problems into four notability tiers:
- Moderately Interesting — solid but narrow results within a subfield
- Solid Result — publishable advances of broader interest
- Major Advance — significant results that would attract wide attention
- Breakthrough — results that would reshape a field
The hardest tier includes problems like finding a faster prime factorization algorithm and adapting Apéry's proof of the irrationality of ζ(3) to other constants. Estimated solving times range from one to four weeks at the low end to three to ten years at the high end. The number of serious human attempts per problem ranges from two or three mathematicians to over fifty.