Stanford and Google's LoT Diffusion Cuts Image Tokens by 4.6x Without Losing Detail

Stanford and Google researchers introduce a diffusion method that spends compute only where detail matters, cutting token counts up to 4.6x.

·
·
·
Stanford and Google's LoT Diffusion Cuts Image Tokens by 4.6x Without Losing DetailPRO
  • Stanford and Google propose LoT Diffusion, allocating tokens by anticipated detail instead of uniformly.
  • Each token is a rectangle of arbitrary size; fine tiles cover subjects, coarse tiles cover backgrounds.
  • Layouts can come from semantic masks, bounding boxes, texture variance, depth, or an agent's importance map.
  • Reports up to 4.6x fewer tokens and 2.5x faster generation at matched quality.
  • Applied to Flux.2 for images and Wan2.1 for video via patch-wise asymmetric flow matching.
  • Beats ToMe-SD, DDiT, and Foveated Diffusion at the same token budget without boundary artifacts.

LoT Diffusion lets detail set the token budget

Image and video diffusion transformers spend as many tokens on a blank wall as they do on a human face. The LoT Diffusion paper from Stanford and Google adapts pretrained models to use small tokens for detailed regions and larger tokens elsewhere. Across the authors’ experiments, the approach reduces token counts by 1.3× to 4.6× while maintaining quality comparable to full-resolution generation.

Uniform grids overpay for blank space

Diffusion transformers, commonly called DiTs, divide an image or video latent into a regular grid of patches. Every patch becomes a token that passes through the transformer at each denoising step, so a flat sky occupying 40% of a frame consumes the same number of tokens as an equally large foreground subject.

Self-attention scales quadratically with sequence length, which makes token count a major cost driver at high resolutions and long video durations. The paper compares LoT with acceleration methods including ToMe-SD, DDiT, and Foveated Diffusion, reporting that those baselines produce more noise or visible boundaries at matched token budgets.

LoT Diffusion examples showing fine tokens over detailed subjects and coarse tokens over simpler image and video regions
LoT covers the full latent canvas with rectangular tokens of different sizes.

The layout becomes the compute budget

LoT accepts a spatial importance map produced before denoising and converts it into a rectangular tiling. Small rectangles reserve more tokens for detailed regions, while large rectangles compress areas that can tolerate lower spatial precision. The paper demonstrates five sources for these maps:

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads