Harvard's Projection Sampling Makes Standard Fine-Tuning Match Reinforcement Learning
Harvard researchers use MCMC sampling to reshape expert data into on-policy traces, letting supervised finetuning match or beat RL on generalization.
- Harvard paper shows SFT can rival RL when training data is reshaped via MCMC sampling.
- Projection sampling uses Metropolis-Hastings to pull expert traces closer to base model distribution in KL.
- Information constraint is folded into the proposer, not used as external verifier, improving efficiency.
- Beats RL and OPSD on chemistry, math, and open-ended tasks with less forgetting.
- Pass@k curves exceed base model, indicating genuinely new capability rather than distribution sharpening.
- More MCMC steps monotonically shrink KL gap and improve downstream finetuning accuracy.
Projection sampling pulls expert traces toward the base model
Harvard researchers Aayush Karan, Sitan Chen, and Yilun Du attribute much of the performance gap between supervised fine-tuning and reinforcement learning to the distribution of training data. Their paper, Finetuning with Sampling: SFT Learns Better Than You Think, introduces a sampling procedure that adapts expert demonstrations to the base model before supervised training begins.
The procedure, called projection sampling, uses Markov chain Monte Carlo to rewrite expert trajectories while preserving required information such as a correct answer or reasoning step. The paper reports that ordinary supervised fine-tuning on these rewritten examples matches or exceeds strong on-policy methods across chemistry, mathematics, and open-ended tasks, with less degradation of the base model’s existing capabilities.
Why raw expert traces cause drift
Supervised fine-tuning, or SFT, trains a model to predict tokens from fixed demonstrations. Those demonstrations can contain solutions that the base model would rarely discover, making SFT useful when successful outputs are scarce or expensive to generate.
Expert demonstrations may also differ sharply from the model’s own output distribution in vocabulary, reasoning structure, and token transitions. Repeatedly fitting those unlikely sequences can move the model away from behaviors learned during pretraining, contributing to catastrophic forgetting.
On-policy reinforcement learning generates candidate trajectories from the current model, scores them with a reward function, and updates the model from successful samples. Its training data therefore stays close to the model’s distribution, but sparse rewards create a bottleneck when the model rarely produces a correct trajectory.
| Method | Training signal | Primary constraint |
|---|---|---|
| Standard SFT |
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.