Synthetic Warm-Up Saves 21B Training Tokens by Boosting Long-Range Retrieval

The largest study of synthetic pre-pretraining shows it saves 21B+ tokens at 3B scale, but the grammatical prior story falls apart.

·
·
Synthetic Warm-Up Saves 21B Training Tokens by Boosting Long-Range RetrievalPRO
  • Largest study of synthetic pre-pretraining to date, covering 500M to 7B parameters and up to 100B pretraining tokens.
  • PPT saves at least 21B pretraining tokens at the 3B scale, making it a cheap net win.
  • Gains hold across web, code, and math mixtures, but vanish when web text is removed.
  • Downstream improvements do not track grammatical acceptability on BLiMP, contradicting the grammatical prior hypothesis.
  • The real mechanism appears to be improved long-range retrieval learned from synthetic sequences.
  • Full paper on arXiv with HTML version here.

Synthetic warm-up gains track long-range retrieval

A broad scaling study finds that pre-pretraining on synthetic sequences improves language-model training through 7B parameters and 100B tokens. Models reach matched downstream performance with fewer main-phase training tokens, including an estimated saving of at least 21B tokens at the 3B scale. The paper finds no consistent support for the proposed grammatical prior; the strongest gains track improved long-range retrieval.

Synthetic sequences before web text

Pre-pretraining, or PPT, adds a short synthetic training phase before conventional pretraining on web text, code, and math. Researchers use generated sequences because they can control the exact structures and dependencies a model must learn.

The canonical task, k-Shuffle Dyck, interleaves several streams of balanced brackets. Solving it requires the model to identify each stream, track nested brackets, and recover relevant symbols across long spans.

Prior work attributed PPT gains to a structural inductive bias. Under that account, exposure to hierarchy and recursion prepares a transformer to learn natural-language syntax more efficiently.

Diagram showing the proposed transfer of a grammatical prior from synthetic pre-pretraining to natural-language training
The study tests whether a grammatical prior explains PPT’s downstream gains.

Earlier tests stopped at 1B

Previous PPT experiments used models no larger than 1B parameters, training budgets below 2B tokens, and data dominated by web text. Those limits left three practical questions unresolved:

  1. Do the gains persist as model size and token budgets increase?
  2. Do they survive mixtures containing substantial code and math?
  3. Which property of a synthetic task predicts useful transfer?

The new grid reaches 100B tokens

The study expands the test across five synthetic tasks, four pretraining mixtures, four model scales, and budgets reaching 100B tokens.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads