MIT Finds a Simple Word Prefix Boosts AI Math Scores by 36 Points

A multi-institution study shows base models can match RL-trained counterparts on math and code simply by prefilling the right opening tokens, with effects traceable to training data.

·
·
·
MIT Finds a Simple Word Prefix Boosts AI Math Scores by 36 PointsPRO
  • Prefilling ".\n\nOkay" lifts Olmo-3-7B MATH-500 pass@1 from 42% to 78% with no weight changes.
  • Prefilling "Alright," lifts Qwen3-14B MATH-500 pass@1 from 72% to 87%, matching its RL variant.
  • KL divergence between base and RL policies peaks at the first two response tokens then flattens.
  • Causal data edits turn arbitrary words like "chicken" into effective reasoning cues.
  • Hidden states after effective cues drift toward training-set reasoning traces.
  • Code and project page: github.com/sophicle/cues, sophielwang.com/cues.

Short response prefixes unlock reasoning in base models

Researchers at MIT, UC Berkeley, the University of Washington, and the Allen Institute for AI report that short response prefixes can recover much of the reasoning performance associated with reinforcement learning. Wang, Dravid, Shao, Farhat, Min, and Efros trace the effect to patterns learned from training data and show that editing those patterns can assign the same behavior to arbitrary words. Their results appear in the paper.

Small cues, large accuracy swings

Reinforcement learning is commonly used to improve language models on math and programming tasks. One method, Group Relative Policy Optimization, or GRPO, samples a group of answers, compares their rewards, and increases the probability of higher-scoring responses. The study examines whether those updates teach new reasoning procedures or increase access to behaviors learned during pretraining.

On MATH-500, prefilling a response with a few characters produced substantial pass@1 gains. Pass@1 measures whether the model’s first generated answer is correct. The notation \n below represents a newline.

Model Response prefix Baseline With prefix Change
Olmo-3-7B .\n\nOkay 42% 78% +36 percentage points
Qwen3-14B Alright, 72% 87% +15 percentage points

No weights changed during these prefill experiments. The researchers inserted the prefix at the start of the model’s response, then allowed generation to continue normally.

Accuracy comparison among base, cue-prefilled, and reinforcement-learned models on math and programming benchmarks
Short prefixes narrow the measured accuracy gap between base and reinforcement-learned models, with results varying by model and benchmark.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads