Pangram Labs Warns AI Web Text Could Triple Pretraining Costs by 2028

Pangram Labs trained 800 language models to measure how AI-generated web text warps pretraining, and built a new scaling law to price each token.

·
·
·
Pangram Labs Warns AI Web Text Could Triple Pretraining Costs by 2028PRO
  • Pangram Labs trained 800 LMs to measure how wild AI-generated web text affects pretraining loss.
  • 31% of August 2026 FineWeb-filtered tokens are AI-generated, up from 10% in June 2024.
  • New scaling law with benefit and harm terms predicts loss 41% better than Chinchilla on larger models.
  • Unfiltered 2026 crawls cost 1.6x the compute of their human subset; forecast 3.0x by end of 2028.
  • FineWeb keeps AI docs 2.3x and DCLM 9.8x more often than human docs, worsening the problem.
  • WildAI corpus, 800 models and code released at github.com/pangramlabs/WildAI under CC BY-NC-SA 4.0.

AI Web Text Changes the Economics of Pretraining

In an August 2026 web crawl, researchers from Pangram Labs and the University of Maryland classified 31% of tokens as AI-generated. A new study then trained 800 language models on controlled mixtures of human and AI web text to measure how each source affects next-token prediction loss. The result depends on the training regime: generated tokens can help data-starved models, but their marginal value saturates and eventually becomes negative as human-data budgets and model sizes grow. Training plans based on the Chinchilla scaling law miss that reversal because they assign equal value to every token.

The team’s estimated AI share rose from roughly 10% in June 2024 to 27.5% in June 2026 and 31% two months later. Common quality filters increased that concentration. FineWeb’s pipeline retained AI-labeled documents at 2.3 times the rate of human-labeled documents, while DCLM’s retained them at 9.8 times the rate. A filtered crawl can therefore contain a higher proportion of generated text than its source crawl.

AI share of web tokens over time, the marginal value of AI tokens, and estimated compute penalties
The study connects the rising share of AI web text with changes in token value and pretraining compute.

Equal-token math reaches its limit

Earlier synthetic-data research largely examined purpose-built training corpora or recursive model-collapse experiments in which models learn from their own outputs. Wild web text has a broader origin. It comes from systems such as GPT-4o, Claude, and Llama, may include human edits, and appears alongside human writing across many topics and formats. The WildAI experiments isolate this material in controlled training mixtures.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads