Looped transformers: three different machines
Classify depth, time, or flow. Nanbeige is the downloadable LLM. RLT lost at 140M. Looped Flows is a 5-7M puzzle net
- Nanbeige4.2-3B is Apache 2.0 on Hugging Face, with
num_hidden_layers: 22andnum_loops: 2, and two passes kept about 75% of a standard Transformer's token efficiency after 28 trillion tokens. - Serving it costs about 2x block compute and a 2x KV cache: 22 layer applications become 44, while the file stays at 4 billion parameters, about 3 billion of them outside the embeddings.
- At about 140 million parameters and 500 million tokens, the Recurrent Looped Transformer scored 3.6884 nats against 3.6420 for a matched Transformer, about 8.4K versus 124K tokens/s on an RTX 5090 (22 versus 1.14 GPU-hours).
- Looped Flows, 5 million parameters on Sudoku and 7 million otherwise, raised Sudoku-Extreme from 74.5% at 8 steps to 97.9% at 128 with the same weights.
- Jakub Pachocki said Astra's computation-graph depth is within a factor of two of GPT-4, and an OpenAI safety researcher wrote that a frontier-scale unmonitorable recurrent model is a day that "is not today."
You've probably heard the buzz around Looped transformers, but which one? Over the past month, that term was applied to three fundamentally different machines.
Before diving into them, let's agree that a LLM writes one token at a time, and a transformer is the layer stack that turns the tokens you already have into the next one. A looped design reuses some of that stack instead of adding more unique layers.
But there are different reuse mechanisms, and if you implement the wrong one, you will either waste a week of training compute or trigger out-of-memory errors by budgeting a 3B-weight model as a standard 3B decode.
Next-token loss is measured in nats, the standard base(e) cross-entropy loss, where lower signifies better prediction. In an independent evaluation at 140M parameters across 500M tokens, the viral Recurrent Looped Transformer underperformed a standard baseline (3.6884 vs. 3.6420 nats) while burning roughly 20x the GPU-hours.
To evaluate any looped model, apply a three-part test: classify its mechanism: depth, time, or flow, verify whether extra inference steps actually buy higher quality, and determine whether its runtime profile belongs in production or gets rejected.
Then we all saw the Astra rumor that made the collision feel urgent. The Information reported that OpenAI's Astra uses "recurrent depth," with a cost and performance upside and a worry that it hides thinking.
Jakub Pachocki, who leads research at OpenAI, answered the same morning: "The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4." Chain-of-thought monitoring, he wrote, is "trending in a negative direction, for reasons not contingent on architecture changes."
An OpenAI safety researcher quote-tweeting him said a frontier-scale recurrent (or otherwise unmonitorable) language model would be one of the darkest in the current AI era, and "this day is not today."
For now, Astra remains unconfirmed speculation: bounded by Pachocki's factor-of-two graph depth constraint, and a safety researcher at OpenAI wrote that this day is not today. I'd keep any loop count out of an Astra runbook.
What follows is the three-way check, the extra-compute test, Nanbeige 4.2's two-pass KV bill, why the time-loop ran at 8.4K tokens/s against 124K on a 5090, and when a 5-7 million flow net is the right sidecar.
What is looped transformer?
A "looped transformer" is one of three machines: extra layer passes on one token, hidden state carried token to token, or a tiny net that denoises a grid from noise.
In early September, Sebastian Raschka drew the initial distinction: reusing layers across depth versus maintaining recurrent state across time. If a README simply advertises a "looped" architecture and you can't categorize which mechanism it uses, you don't actually know what you are implementing.
A concrete check:
Don't miss what's next in AI
Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.
- Full access to in-depth AI research breakdowns
- Be the first to know what's trending before it hits mainstream
- Daily curated papers, repos, and industry moves