Sharpening Tax Shows How RL Post-Training Quietly Kills AI Agent Retries

A new study finds reinforcement learning post-training trades solution coverage for consistency on agentic tasks, and proposes an adaptive sampler that avoids the hit.

·
·
·
Sharpening Tax Shows How RL Post-Training Quietly Kills AI Agent RetriesPRO
Read2 min
TypePaper
SubtopicRlhf · Tool Use · Multi Agent
  • New paper introduces Sharpening Tax, measuring coverage lost when RL post-training sharpens LLM behavior on agentic tasks.
  • Base models with a light harness often beat post-trained twins at pass@K despite worse pass@1 accuracy.
  • Post-training bimodalizes outcomes: more always-solved and never-solved tasks, fewer tasks where retrying helps.
  • Tax is positive in 36 of 42 model-benchmark cases and predictable from only 8 rollouts.
  • PTGS adapts sampling temperature per prompt via Beta posterior, improving both pass@1 and pass@128.
  • Code at github.com/changdaeoh/sharpening-tax; paper at arxiv.org/abs/2610.01509.

Reinforcement-learning post-training may improve an agent’s first attempt while reducing the range of tasks it can solve across repeated attempts. A new preprint reports this pattern in multi-turn, tool-using benchmarks: post-training concentrates probability on behaviors already available in the base model, raising pass@1 while often lowering coverage at larger sampling budgets.

The authors call this reduction the Sharpening Tax. They compare 14 base and post-trained model pairs from four families on three benchmarks, producing 42 model-benchmark evaluations. The tax appears in most settings, can be estimated from a small rollout sample, and correlates with related metrics.

Systems that generate several trajectories and select one with a verifier depend on broad solution coverage. A checkpoint that performs better on its first attempt can still reach fewer tasks after many retries, limiting the returns from best-of-N sampling.

Agents raise the bar

Earlier evidence for the sharpening hypothesis came mainly from math and code tasks. In those settings, reinforcement learning appears to redistribute probability toward successful outputs the pretrained model could already produce, improving consistency without expanding the observed solution set.

Agent benchmarks provide a stronger test because they require models to call tools, interpret observations, recover from errors, and plan across multiple turns. Finding the same trade-off there suggests that post-training does not reliably add those behaviors within the environments studied.

Base models catch up on retries

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads