MIT Finds a Simple Word Prefix Boosts AI Math Scores by 36 Points
A multi-institution study shows base models can match RL-trained counterparts on math and code simply by prefilling the right opening tokens, with effects traceable to training data.
- Prefilling ".\n\nOkay" lifts Olmo-3-7B MATH-500 pass@1 from 42% to 78% with no weight changes.
- Prefilling "Alright," lifts Qwen3-14B MATH-500 pass@1 from 72% to 87%, matching its RL variant.
- KL divergence between base and RL policies peaks at the first two response tokens then flattens.
- Causal data edits turn arbitrary words like "chicken" into effective reasoning cues.
- Hidden states after effective cues drift toward training-set reasoning traces.
- Code and project page: github.com/sophicle/cues, sophielwang.com/cues.
Short response prefixes unlock reasoning in base models
Researchers at MIT, UC Berkeley, the University of Washington, and the Allen Institute for AI report that short response prefixes can recover much of the reasoning performance associated with reinforcement learning. Wang, Dravid, Shao, Farhat, Min, and Efros trace the effect to patterns learned from training data and show that editing those patterns can assign the same behavior to arbitrary words. Their results appear in the paper.
Small cues, large accuracy swings
Reinforcement learning is commonly used to improve language models on math and programming tasks. One method, Group Relative Policy Optimization, or GRPO, samples a group of answers, compares their rewards, and increases the probability of higher-scoring responses. The study examines whether those updates teach new reasoning procedures or increase access to behaviors learned during pretraining.
On MATH-500, prefilling a response with a few characters produced substantial pass@1 gains. Pass@1 measures whether the model’s first generated answer is correct. The notation \n below represents a newline.
| Model | Response prefix | Baseline | With prefix | Change |
|---|---|---|---|---|
| Olmo-3-7B | .\n\nOkay |
42% | 78% | +36 percentage points |
| Qwen3-14B | Alright, |
72% | 87% | +15 percentage points |
No weights changed during these prefill experiments. The researchers inserted the prefix at the start of the model’s response, then allowed generation to continue normally.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.