NVIDIA's PivotOPD Teaches AI Agents to Fix Early Mistakes Before They Spiral
NVIDIA's PivotOPD pinpoints the single early mistake that dooms most agent rollouts, then teaches the student to both avoid it and recover.
- NVIDIA's PivotOPD targets the single early mistake that causes most multi-turn agent failures.
- Over half of failed Qwen3 rollouts hinge on one early pivotal mistake.
- Correcting that turn lifts replayed success from 8.2% to 59% in Qwen3-8B tests.
- Combines preventive reverse-KL distillation with forward-KL recovery distillation on post-mistake states.
- Beats 13 baselines on ALFWorld, WebShop, Search QA; +5.5% ALFWorld with Qwen3-1.7B.
- Transfers across families: +3.2% resolve rate for Nemotron-3.5 on SWE-Bench Verified.
PivotOPD trains agents where failure begins
NVIDIA Research has proposed PivotOPD, a variant of on-policy distillation for multi-turn language agents. The method focuses training on an early decision that derails a failed rollout, then teaches the model to avoid that action and recover from the resulting state. Tests on embodied tasks, web navigation, search-based question answering, and software engineering report gains over standard OPD and other baselines.
How one turn derails a rollout
The authors define a pivotal mistake as an action that moves an agent farther from completing its task. Across failed rollouts from three Qwen3 models ranging from 8 billion to 235 billion parameters, more than half contained such a mistake, usually near the beginning of the interaction.
The clearest intervention study examined 72 failed Qwen3-8B trajectories with oracle-labeled pivotal turns. Replacing the pivotal action during replay raised success from 8.2% to 59.0%. Keeping the original mistake and supplying correct actions for the next few turns produced nearly the same recovery rate, suggesting that a short training window can still rescue the trajectory.
Why ordinary OPD misses the exit
On-policy distillation lets a student model generate its own trajectories while a teacher provides target probabilities for each generated token. The student therefore trains on the states it actually reaches, including states created by its own mistakes.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.