QuadTok Cuts Image Tokens by 10% by Spending Them Where Detail Lives

A hierarchical quadtree tokenizer spends more tokens on busy image regions and fewer on flat ones, cutting token counts about 10% while hitting 2.08 gFID on ImageNet.

·
·
·
QuadTok Cuts Image Tokens by 10% by Spending Them Where Detail LivesPRO
  • QuadTok replaces fixed grid tokenization with a quadtree that refines only visually complex regions.
  • Saves ~10% tokens on ImageNet and 9% zero-shot on COCO versus a 256-token grid baseline.
  • 947M GPT-style generator reaches 2.08 gFID on ImageNet 256x256 with a kinship causal mask.
  • Two-level tokenizer: 64 coarse tokens plus optional fine refinements, 64-320 tokens total per image.
  • Quadtree topology conditioning enables zero-shot spatial layout control without ControlNet-style modules.
  • Tokenizer code and weights released under Apache-2.0 at github.com/myc634/QuadTok; generator not yet public.

QuadTok spends image tokens where detail lives

Most image tokenizers divide a picture into equal-sized patches and pass the resulting sequence to a transformer. Flat sky receives the same token budget as a detailed face. QuadTok, developed by Matthew Mao and collaborators, replaces that fixed grid with a quadtree that adds tokens only where finer reconstruction improves perceptual quality.

Reported results at 256 × 256 resolution
Benchmark Result Context
ImageNet reconstruction About 10% fewer tokens Roughly 230 tokens on average versus a fixed 256-token grid, at comparable reconstruction fidelity
COCO reconstruction About 9% fewer tokens Zero-shot evaluation using the ImageNet-trained tokenizer, without COCO fine-tuning
ImageNet generation 2.08 gFID A 947M-parameter GPT-style generator; lower gFID indicates closer agreement with the real-image distribution

The Apache-2.0 release on GitHub contains tokenizer code and weights. The 947M-parameter generator checkpoint remains unpublished, leaving the reported generation result unavailable through official weights.

Detail gets the budget

Vector-quantized tokenizers used by models such as VAR, LlamaGen, and Parti convert image regions into discrete codebook entries that a transformer can process like vocabulary tokens. A common configuration maps a 256 × 256 image to a 16 × 16 grid, producing 256 tokens. During training, standard transformer attention grows with the square of sequence length; during autoregressive sampling, every additional token adds another sequential decoding step.

Compressed tokenizers such as TiTok and One-D-Piece produce shorter one-dimensional sequences. Their token positions have a weaker correspondence to specific image regions, which complicates coordinate-based conditioning and editing. QuadTok retains an explicit two-dimensional hierarchy while allowing each image to use a different number of tokens.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads