UC Berkeley's EasyPPO Stops AI Training Collapses With Three Critic Fixes

A trio of surgical tweaks to PPO's critic eliminates training collapse in LLM reinforcement learning, delivering up to 14.89% score gains without touching the actor.

·
·
UC Berkeley's EasyPPO Stops AI Training Collapses With Three Critic FixesPRO
  • EasyPPO stabilizes PPO for LLM training with three critic-side fixes, leaving the actor update untouched.
  • Actor-only overlong filtering trains the critic on truncated rollouts so it learns truncation is costly.
  • Noise-normalized critic loss prevents high-variance prompts from dominating gradient updates in finite batches.
  • Smaller critic mini-batches with per-step clipping confine outlier rollout influence during optimization.
  • Relative gains over PPO: 14.89% on FrontierCS coding, 2.28% on AIME24, 9.47% on Search-R1.
  • Code and Qwen3.5-9B configs available at github.com/EasyPPO/EasyPPO.

EasyPPO traces PPO collapses to the critic

Researchers at UC Berkeley and collaborating institutions report in the EasyPPO paper that unstable value learning drives PPO failures across three large language model reinforcement learning settings. Their method stabilizes the critic with three targeted changes while preserving PPO’s policy objective and actor optimizer.

PPO remains common in LLM post-training because it can learn from continuous, binary, and multi-step rewards. Its reliability depends on a critic whose predictions shape every policy update, so critic errors can stall training or erase gains from an otherwise healthy run. EasyPPO addresses that failure point without introducing a new reinforcement learning algorithm.

How critic errors reach the policy

PPO trains an actor alongside a critic. The actor generates responses, while the critic estimates the expected future return, or cumulative reward, for each prompt and partial response. PPO combines observed returns with those estimates to calculate advantages, which determine how strongly each generated token should be reinforced.

Inaccurate value estimates produce noisy advantages. The actor then strengthens or weakens tokens using a distorted learning signal, allowing critic instability to spread into the policy even when the policy loss and optimizer behave as designed.

Two failure modes feed collapse

Truncation disappears from training

A filtering practice common in GRPO implementations drops rollouts that reach the token limit before completion. Applying that filter to both the actor and critic conditions training on completed responses: the critic never learns that exhausting the context budget is undesirable, and the reported reward excludes the failed, truncated trajectories.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads