SenseTime's Looped-DiT Beats a 1.7B Model Using Just 260M Parameters
A 260M-parameter text-to-image model beats one 6.5x larger by running the same transformer blocks multiple times per denoising step.
- Looped-DiT runs shared transformer blocks multiple times per denoising step, scaling compute without scaling parameters.
- A 260M model beats a 6.5x larger baseline using 4.9x less inference compute on text-to-image benchmarks.
- Naive looping degrades quality; the paper traces it to weak intermediate supervision and runaway attention updates.
- Fixes: deep supervision at every loop plus self-modulating attention (XSA) to protect local spatial information.
- Extra loops beat extra denoising steps under matched FLOPs, and later loops visibly correct earlier mistakes.
- Code is available at github.com/OpenSenseNova/Looped-DiT; paper here.
Looped-DiT turns repeated depth into a diffusion scaling axis
Text-to-image diffusion models usually scale through larger networks or additional denoising steps. Researchers from SenseTime Research, Tsinghua University, and Nanyang Technological University propose another option in the Looped-DiT paper: reuse the same transformer blocks several times within each denoising step. The model gains computational depth at inference without adding parameters.
At every denoising timestep, a conventional diffusion transformer processes the current noisy image representation once before advancing the sampler. Looped-DiT applies its shared block stack repeatedly before moving to the next timestep. Loop count becomes a configurable scaling control alongside model size and sampler length.
| Measure | Reported result |
|---|---|
| Parameters | 260 million, compared with about 1.7 billion for the larger baseline |
| Image quality | The 260M model surpasses the 6.5-times-larger model across multiple text-to-image benchmarks |
| Inference compute | About one-fifth of the larger model’s compute, reported as a 4.9-times reduction |
A smaller weight set reduces model storage and memory bandwidth requirements, while repeated execution spends additional compute on the same parameters. The trade is especially relevant when model weights exceed available VRAM but longer inference remains acceptable.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.