Epoch Patches EBR-Bench After GPT-6 Astra Breaks It With Unlimited Turns

Epoch AI bans a game-breaking card from its long-horizon board game benchmark after GPT-6 Astra aced it, and tests whether multi-agent setups help.

·
·
·
Epoch Patches EBR-Bench After GPT-6 Astra Breaks It With Unlimited Turns
Read5 min
TypeNews
SubtopicMulti Agent
  • Epoch AI updated EBR-bench, its long-horizon board game benchmark built on Earthborne Rangers.
  • GPT-6 Astra hit a perfect score on over half of attempts by exploiting one overpowered card.
  • That card bypasses the game's fatigue mechanic and is now banned in the default setting.
  • Even with the ban, Astra averages a 50% jump over prior top models, scoring up to 20 of 21.
  • Multi-agent scaffolds via Inspect Deep Agent improved deck exploration but did not raise scores.
  • Epoch expects saturation within months and will move to harder, more diverse games next.

Epoch patches EBR-bench after Astra finds an unlimited-turn combo

Epoch AI has released what it calls the final protocol update for EBR-bench, its long-horizon benchmark based on the cooperative card game Earthborne Rangers. The revised rules ban a card that OpenAI’s GPT-6 Astra used to achieve frequent perfect scores and add an optional multi-agent scaffold for testing groups of model copies.

The changes address a benchmark approaching saturation, where leading models score near the maximum and the leaderboard loses its ability to distinguish between them.

  • Card ban: Removes an unlimited-turn combination that bypasses the benchmark’s main tactical constraint.
  • Multi-agent option: Lets one model delegate work to as many as four copies of itself.
  • Smaller evaluations: Reduces future experiments from 10 samples to five and retires several alternate run configurations.

Why a board game works here

Earthborne Rangers gives models a long sequence of connected decisions rather than a puzzle with one predetermined solution. A playthrough takes human players roughly two to four hours, requiring them to build decks, manage resources, adapt tactics, and pursue objectives across many turns. These characteristics make the game useful for measuring long-horizon performance on complex, multi-step tasks.

Epoch’s original study found that models struggled to explore varied decks, improve through repeated play, and balance immediate tactical choices against longer-term goals.

GPT-6 Astra changed the leaderboard. It achieved a perfect score in more than half of its initial attempts and averaged roughly 19.8 out of 21. Claude Opus 5, the next-highest model, scored 10.5.

Bar chart comparing EBR-bench scores across models, led by GPT-6 Astra
GPT-6 Astra approached the benchmark’s 21-point ceiling before the card ban.

One card erased the turn limit

Epoch traced Astra’s dominant performance to one card. The strategy follows the game’s rules, and the strongest human baseline player also used it to reach the maximum score. Combined with several supporting cards, however, it provides an effectively unlimited number of turns and bypasses the tactical pressure that EBR-bench was designed to measure.

A run can complete as many as 21 objectives within a limited turn budget. Obstacles and enemies impose fatigue, which reduces the turns available for reaching those objectives. Ordinary decks cannot replenish turns indefinitely. The banned combination removes that constraint, making fatigue management largely irrelevant.

Selected results behind the protocol change
Measure Result
Astra’s average before the ban About 19.8 out of 21
Astra’s best score after the ban 20 out of 21
Astra’s turns with the card 88
Astra’s turns after the ban 42
Top human’s longest combo run 183 turns
Chart comparing turn counts when the banned card was allowed and removed
The card combination more than doubled Astra’s turn count and supported a 183-turn human run.

Astra learned to use the combination quickly. It scored 11 out of 21 on its first attempt and 21 on its second. The top human baseline required six playthroughs to reach 21. Epoch avoids labeling the result as superhuman reasoning, but the run shows that Astra can absorb the rules and card pool quickly enough to identify and apply a game-breaking strategy.

The leaderboard survives the patch

Changing a benchmark after publication complicates comparisons, so Epoch checked whether earlier models depended on the same card. Three recent models included it in their decks to varying degrees, but none used the unlimited-turn combination, and removing the card had no statistically significant effect on their topline scores.

Epoch will report card-banned results for Claude Fable 5.1, Claude Opus 5, GPT-5.6 Sol, GPT-6 Astra, and future entrants. Historical results for other models remain based on the earlier protocol.

Astra retains a substantial lead under the revised rules. Its best card-banned score is 20 out of 21, and its average remains roughly 50% higher than those of the previous strongest models. Its fatigue management remains mediocre, however, with no measurable improvement across repeated attempts. That gap shows how a near-ceiling objective score can coexist with weaker tactical play.

Subagents broaden deck selection

Epoch’s second experiment tested whether model copies could improve performance through parallel exploration. The setup uses the Inspect Deep Agent harness from the UK AI Security Institute, a software wrapper that lets a model plan, keep notes, and delegate tasks to helper copies of itself.

Each system could create as many as four subagents running EBR-bench instances in parallel. The full system shared the same 10-playthrough budget used for single-agent evaluations, preventing additional game attempts from becoming an independent advantage.

  • Multi-agent systems explored a wider range of decks for most of the four tested models.
  • Only Claude Opus 5 showed a statistically significant increase in deck diversity.
  • The additional exploration produced no meaningful change in topline scores.
  • The score effect varied in direction across models.
  • Astra already explored many deck types as a single agent.

The experiment found no automatic performance gain from adding model copies. Lower-scoring models tried more strategies with subagents but failed to convert that variety into better outcomes. Astra’s single-agent behavior already supplied much of the exploration that the scaffold was intended to encourage.

Near-perfect scores narrow the signal

EBR-bench was designed to resemble open-ended discovery tasks, where an agent must explore a large option space and improve through experience. Astra’s performance demonstrates how quickly a capable model can maximize a benchmark by finding a legal strategy that bypasses its intended challenge.

The result also exposes a measurement problem. EBR-bench’s objective score records how much a model accomplishes, while fatigue and turn counts reveal how it plays. Reporting those measures together helps distinguish tactical improvement from reliance on a degenerate strategy.

Epoch expects EBR-bench to reach full saturation within months and plans to move to more complex games. Its technical update also retires 18-playthrough evaluations and dual compaction-threshold runs, which tested different points for summarizing an agent’s accumulated context. Future experiments will use five samples instead of 10, reducing evaluation cost while increasing the uncertainty around average scores.

Once several models cluster near 21 points, EBR-bench will offer limited separation at the top of the leaderboard. More complex environments will need to preserve long-term planning demands while resisting shortcuts that erase the constraints under evaluation.

Trending
  • No trending articles

Comments

avatar

Next Reads