GPT-6 Astra Cracks a Decade-Old Voting Theory Problem Nobody Could Solve

Epoch AI marks the first Major Advance on its unsolved-math benchmark, with GPT-6 Astra driving the proof in an interactive session with three mathematicians.

·
·
GPT-6 Astra Cracks a Decade-Old Voting Theory Problem Nobody Could Solve
  • Epoch AI marked The Core in Approval-Based Committee Elections as solved, the first Major Advance on the board.
  • Becker, Greger, and Peters credit GPT-6 Astra for the primary idea in a lengthy interactive session.
  • The proof shows the core is always non-empty, so no counterexample exists to the 2017 Aziz et al. question.
  • Epoch's new human + AI label reflects active human elicitation with core ideas coming from the model.
  • Board status: 1 of 6 Major Advance, 2 of 18 Solid Result, 5 of 22 Moderately Interesting, 0 of 3 Breakthrough solved.
  • Context: Astra saturated FrontierMath Tier 4 at 98% but only 2 of 68 on FrontierMath Erdős.

GPT-6 Astra helps close a decade-old voting theory problem

Epoch AI has marked a committee-election problem as solved after researchers proved that its requested counterexample cannot exist. “The Core in Approval-Based Committee Elections” is the first solved problem in the Major Advance category of FrontierMath: Open Problems.

The proof emerged from a lengthy exchange between GPT-6 Astra and Becker, Greger, and Dominik Peters, according to Epoch. The resulting preprint credits the model with the primary idea. Epoch classifies the work as an interactive human-and-AI collaboration conducted with larger compute budgets and varied agent setups, outside the constraints of a one-shot benchmark run.

The counterexample that cannot exist

In an approval-based multiwinner election, each voter selects the candidates they approve, and the election chooses a committee of k members. A voter’s utility is the number of approved candidates on that committee.

Aziz, Brill, Conitzer, Elkind, Freeman, and Walsh posed the open question in 2017. They asked whether an election could have an empty core. A committee belongs to the core when no sufficiently large coalition can propose an alternative slate T that every coalition member strictly prefers. A group large enough to claim |T| seats must contain at least |T|n/k of the election’s n voters.

Becker, Greger, and Peters proved that the core is always non-empty. The requested election therefore does not exist. Epoch accepted the theorem as a solution because its credit policy counts any resolution of the underlying mathematical question, including a proof that the requested object is impossible.

Credit follows the workflow

Epoch formerly reserved its AI-solved designation for autonomous results. It added a human-and-AI category after researchers began reporting proofs that depended on iterative model guidance and would probably have remained undiscovered without it.

Peters, who proposed the problem for the benchmark, told Epoch that the proof required both Astra’s contribution and sustained direction from the researchers. The model could not produce the result from a bare prompt. Epoch recommends counting this category as AI solutions when an analysis requires a two-way classification.

Epoch applied the same attribution standard to an inverse Galois problem involving the Mathieu group M23. It withheld AI-solved status because the authors could not separate the model’s reasoning from the humans’ work, and key decisions came from the mathematicians’ judgment.

A first for the major tier

FrontierMath: Open Problems collects unsolved research questions and pairs them with custom programs that can check candidate outputs. This case required a mathematical proof because the requested counterexample cannot be submitted to the original verifier. Epoch evaluated the theorem under its broader resolution policy.

The board expanded to 50 problems and currently lists 49 active entries after Epoch removed one question. Its published tally is:

FrontierMath: Open Problems by notability tier
Tier Solved Active
Moderately interesting 5 22
Solid result 2 18
Major advance 1 6
Breakthrough 0 3

Earlier AI-credited solutions cluster in the two lower tiers. They include an Anthropic team’s Hadamard construction of order 668 using Claude, rational-point constructions on genus-2 curves, and short superpermutations over 8, 9, and 10 symbols.

Strong scores carry steep costs

Astra has also raised the reported leading score on Epoch’s separate FrontierMath Tier 4 suite to 98%. The top score climbed from 5% to 98% during roughly 14 months following the suite’s July 2025 launch.

Results on the harder FrontierMath Erdős benchmark remain sparse. Astra solved 2 of its 68 historically unsolved problems during the official run and 5 across all attempts. Every other tested model, including GPT-5.6 Sol and Claude Fable 5.1, solved none.

Epoch’s published figures show that one Astra counterexample required 15 hours and $218 in compute. Producing five Erdős solutions across all attempts reportedly consumed more than $220,000, illustrating how exploratory research runs can differ from fixed-budget benchmark scores.

Access and ground truth remain unsettled

OpenAI funded part of FrontierMath’s development and has exclusive access to a portion of the benchmark. Model comparisons therefore do not all occur with identical access to the underlying problem set.

Epoch also removed a problem about stretched Littlewood-Richardson coefficients after losing confidence that a counterexample small enough for its verifier exists. That removal reduced the active board from 50 to 49 problems and shows that the benchmark’s questions and expected answers can change as their mathematics receives further scrutiny.

The clearest signal is collaboration

The result combines a proof of a question open since 2017, a public preprint by named mathematicians, formal recognition from the benchmark operator, and explicit attribution of the central idea to a language model. Its Major Advance classification places it above the construction problems that account for most earlier AI-credited solves.

The documented capability is interactive research support, with autonomous performance on a bare prompt still unestablished. Astra supplied an idea that helped resolve the problem; the mathematicians directed the search, developed the argument, and produced the proof.

Trending
  • No trending articles

Comments

avatar

Next Reads