Pangram Labs Warns AI Web Text Could Triple Pretraining Costs by 2028
Pangram Labs trained 800 language models to measure how AI-generated web text warps pretraining, and built a new scaling law to price each token.
- Pangram Labs trained 800 LMs to measure how wild AI-generated web text affects pretraining loss.
- 31% of August 2026 FineWeb-filtered tokens are AI-generated, up from 10% in June 2024.
- New scaling law with benefit and harm terms predicts loss 41% better than Chinchilla on larger models.
- Unfiltered 2026 crawls cost 1.6x the compute of their human subset; forecast 3.0x by end of 2028.
- FineWeb keeps AI docs 2.3x and DCLM 9.8x more often than human docs, worsening the problem.
- WildAI corpus, 800 models and code released at github.com/pangramlabs/WildAI under CC BY-NC-SA 4.0.
AI Web Text Changes the Economics of Pretraining
In an August 2026 web crawl, researchers from Pangram Labs and the University of Maryland classified 31% of tokens as AI-generated. A new study then trained 800 language models on controlled mixtures of human and AI web text to measure how each source affects next-token prediction loss. The result depends on the training regime: generated tokens can help data-starved models, but their marginal value saturates and eventually becomes negative as human-data budgets and model sizes grow. Training plans based on the Chinchilla scaling law miss that reversal because they assign equal value to every token.
The team’s estimated AI share rose from roughly 10% in June 2024 to 27.5% in June 2026 and 31% two months later. Common quality filters increased that concentration. FineWeb’s pipeline retained AI-labeled documents at 2.3 times the rate of human-labeled documents, while DCLM’s retained them at 9.8 times the rate. A filtered crawl can therefore contain a higher proportion of generated text than its source crawl.
Equal-token math reaches its limit
Earlier synthetic-data research largely examined purpose-built training corpora or recursive model-collapse experiments in which models learn from their own outputs. Wild web text has a broader origin. It comes from systems such as GPT-4o, Claude, and Llama, may include human edits, and appears alongside human writing across many topics and formats. The WildAI experiments isolate this material in controlled training mixtures.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.