AC2 Beats GRPO on Math Proofs Using 2.5x Fewer Decoding FLOPs

A new actor-critic recipe called AC2 trains LLMs on long reasoning tasks without rolling every trajectory to completion, hitting GRPO's peak with 2.5x fewer decoding FLOPs.

·
·
·
AC2 Beats GRPO on Math Proofs Using 2.5x Fewer Decoding FLOPsPRO
  • AC2 trains LLMs with RL using partial rollouts scored by a learned critic, not just terminal rewards.
  • Beats GRPO on IMO-ProofBench with 2.5x fewer decoding FLOPs training Qwen3-4B.
  • Credit is assigned over 10,000-token action chunks instead of individual tokens or whole trajectories.
  • Local readiness gates critic use per-problem until its error on that problem is small.
  • Reference solutions from past successful rollouts are injected into the critic's prompt to boost accuracy.
  • Policy and critic share weights; the critic value is read via a natural-language prompt.

AC2 trims long-rollout RL with a trusted critic

Reinforcement learning with verifiable rewards usually runs each language-model response to completion because the verifier scores only the final output. Long proofs can consume tens of thousands of generated tokens before producing one reward, and common policy-gradient methods assign that trajectory-level signal across every token. The paper Trust the Critic More proposes stopping selected rollouts early when a learned value function can reliably score an intermediate state.

In the paper’s main experiment, Actor-Critic with Action Chunking (AC2) trains Qwen3-4B on proof generation and exceeds GRPO’s peak validation score of 18.5% on IMO-ProofBench. The authors report reaching that result with about 40% of GRPO’s decoding FLOPs, a 2.5-fold reduction. They also released accompanying code.

Reported setup and result
Component Details
Model Qwen3-4B
Training data FineProofs-RL
Evaluation IMO-ProofBench validation set
Baseline GRPO
Action chunk 10,000 tokens
Reported result Above 18.5% with 2.5-fold fewer decoding FLOPs

Why long proofs expose GRPO

GRPO samples groups of responses, scores their final outputs, and derives relative advantages from those scores. In proof generation, a verifier may provide a useful reward only after the complete argument has been generated. Every token in a response then receives the same trajectory-level advantage, even when one intermediate step caused the eventual failure. This coarse credit assignment also prevents training from using an unfinished trajectory.

Actor-critic methods add a value model, or critic, that estimates the expected final return from an intermediate state. A sufficiently accurate critic can provide earlier feedback and reduce the need for complete rollouts. Prediction errors can bias policy updates, however, which has led many language-model RL systems to use critics only as variance-reducing baselines while still collecting terminal rewards.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads