Meta FAIR's Agentic Meta-Reasoning Beats Codex and Claude Code on Long AI Tasks
A controller-worker split lets agents plan which partial work to extend, discard, or ship, pushing ProgramBench scores well past Codex and Claude Code.
- New inference harness separates a controller from workers for long-horizon agent runs, detailed in the arXiv paper.
- Hits 71.5% on ProgramBench with GPT-5.5 versus 58.0% for Codex on the same task.
- With Opus 4.8 reaches 67.2% against 65.5% for Claude Code on ProgramBench.
- Gains 3.6 to 4.2 points over direct-control baselines across reasoning and proof benchmarks.
- Keeps improving as compute budget grows, where simpler scaffolds plateau.
- Controller overhead actively hurts on tiny budgets, so not for short runs.
Long-running agents get a dedicated control loop
Researchers from Meta FAIR, Princeton, and Carnegie Mellon University propose agentic meta-reasoning, a harness that gives long-running AI agents a separate loop for deciding what to do next. In the paper, the method outperforms direct-control baselines across four task categories and records higher ProgramBench scores than configurations using Codex and Claude Code. The practical claim is specific: additional inference compute becomes more useful when a controller summarizes progress, preserves intermediate artifacts, and allocates work under a fixed budget.
Why control degrades over long runs
Long-horizon agents make many sequential model calls, with each decision depending on earlier attempts, tool results, and partial solutions. Most agent harnesses place planning and execution in the same expanding context window. As the transcript grows, the model must recover what succeeded, what failed, and which branch deserves more work before it can take the next step.
That structure creates a control problem. Sampling more attempts adds useful compute, but it also adds history that the agent must interpret. A weak summary can lead the model to repeat failed work, abandon promising artifacts, or spend its remaining budget on branches with little chance of success.
Control moves outside the transcript
The proposed architecture assigns task execution to workers and run management to a controller. The controller receives a compact state, inspects stored artifacts, accounts for the remaining compute budget, and chooses the next action. Raw worker transcripts stay outside its context, reducing the amount of history it must process at each decision.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.