Pleias Trains Small Open Models on SYNTH Using 140x Fewer Tokens

Pleias unveils SYNTH, an open fully synthetic Wikipedia-derived corpus, plus Baguettotron and Monad models that match larger baselines with 10-140x less data.

·
·
Pleias Trains Small Open Models on SYNTH Using 140x Fewer TokensPRO
Read2 min
TypePaper
  • Pleias released SYNTH, the first open fully-synthetic pretraining corpus, derived from 58,698 Wikipedia articles.
  • Baguettotron (321M, 80 layers) hits SOTA for sub-400M models on MMLU, GSM8K, and HotPotQA.
  • Monad is a 56M transformer with above-random MMLU performance, the smallest viable generalist LM.
  • Training collapses pre, mid, and post-training into a single stage with reasoning traces built in.
  • Models reach competitive accuracy with 10 to 140x fewer training tokens than peers.
  • Dataset is CC-BY-4.0, models are Apache 2.0, trained on 16 H100s via Jean Zay.

Pleias, a French-German AI lab, has published a paper describing SYNTH, an open synthetic corpus designed to combine factual learning, reasoning, and instruction following in one model-training stage. The lab trained several small language models exclusively on the corpus and reports competitive results with substantially fewer token exposures than comparable web-trained models.

The approach targets a persistent problem in open model development. Leading labs augment web data with proprietary reasoning traces, instructions, and other synthetic examples, but rarely publish those mixtures or explain how each component affects learning. SYNTH provides a corpus, model weights, and a documented generation recipe that other teams can inspect and adapt.

Where the single stage begins

Conventional language-model pipelines often separate pre-training, continued pre-training, instruction tuning, and reinforcement learning. Each phase teaches a different mix of knowledge, behavior, and task-specific skills.

Pleias folds those objectives into the examples used for initial model training. Each sample can combine source-grounded knowledge, an instruction, a response, and a reasoning trace. Teams shipping assistants may still need later work on safety, preference alignment, tool use, or a specific conversational style.

The term “single-stage” applies to training the final language model. Building SYNTH remains a two-step process in which auxiliary models first learn to generate structured examples and then produce the full corpus at scale.

Wikipedia, multiplied

SYNTH begins with 58,698 selected Wikipedia articles obtained through Wikimedia Enterprise’s Structured Wikipedia dataset. Pleias expands that source material into 79,648,272 samples containing more than 41 billion words, or about 75 billion tokens with the lab’s tokenizer.

The pipeline uses knowledge back-translation: a generator converts a source passage into prompts, answers, and reasoning steps intended to reconstruct or apply the information in that passage. This preserves a link between synthetic examples and identifiable source material, although grounding cannot guarantee that every generated statement is correct.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads