AC2 Beats GRPO on Math Proofs Using 2.5x Fewer Decoding FLOPs
A new actor-critic recipe called AC2 trains LLMs on long reasoning tasks without rolling every trajectory to completion, hitting GRPO's peak with 2.5x fewer decoding FLOPs.
- AC2 trains LLMs with RL using partial rollouts scored by a learned critic, not just terminal rewards.
- Beats GRPO on IMO-ProofBench with 2.5x fewer decoding FLOPs training Qwen3-4B.
- Credit is assigned over 10,000-token action chunks instead of individual tokens or whole trajectories.
- Local readiness gates critic use per-problem until its error on that problem is small.
- Reference solutions from past successful rollouts are injected into the critic's prompt to boost accuracy.
- Policy and critic share weights; the critic value is read via a natural-language prompt.
AC2 trims long-rollout RL with a trusted critic
Reinforcement learning with verifiable rewards usually runs each language-model response to completion because the verifier scores only the final output. Long proofs can consume tens of thousands of generated tokens before producing one reward, and common policy-gradient methods assign that trajectory-level signal across every token. The paper Trust the Critic More proposes stopping selected rollouts early when a learned value function can reliably score an intermediate state.
In the paper’s main experiment, Actor-Critic with Action Chunking (AC2) trains Qwen3-4B on proof generation and exceeds GRPO’s peak validation score of 18.5% on IMO-ProofBench. The authors report reaching that result with about 40% of GRPO’s decoding FLOPs, a 2.5-fold reduction. They also released accompanying code.
| Component | Details |
|---|---|
| Model | Qwen3-4B |
| Training data | FineProofs-RL |
| Evaluation | IMO-ProofBench validation set |
| Baseline | GRPO |
| Action chunk | 10,000 tokens |
| Reported result | Above 18.5% with 2.5-fold fewer decoding FLOPs |
Why long proofs expose GRPO
GRPO samples groups of responses, scores their final outputs, and derives relative advantages from those scores. In proof generation, a verifier may provide a useful reward only after the complete argument has been generated. Every token in a response then receives the same trajectory-level advantage, even when one intermediate step caused the eventual failure. This coarse credit assignment also prevents training from using an unfinished trajectory.
Actor-critic methods add a value model, or critic, that estimates the expected final return from an intermediate state. A sufficiently accurate critic can provide earlier feedback and reduce the need for complete rollouts. Prediction errors can bias policy updates, however, which has led many language-model RL systems to use critics only as variance-reducing baselines while still collecting terminal rewards.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.