QuadTok Cuts Image Tokens by 10% by Spending Them Where Detail Lives
A hierarchical quadtree tokenizer spends more tokens on busy image regions and fewer on flat ones, cutting token counts about 10% while hitting 2.08 gFID on ImageNet.
- QuadTok replaces fixed grid tokenization with a quadtree that refines only visually complex regions.
- Saves ~10% tokens on ImageNet and 9% zero-shot on COCO versus a 256-token grid baseline.
- 947M GPT-style generator reaches 2.08 gFID on ImageNet 256x256 with a kinship causal mask.
- Two-level tokenizer: 64 coarse tokens plus optional fine refinements, 64-320 tokens total per image.
- Quadtree topology conditioning enables zero-shot spatial layout control without ControlNet-style modules.
- Tokenizer code and weights released under Apache-2.0 at github.com/myc634/QuadTok; generator not yet public.
QuadTok spends image tokens where detail lives
Most image tokenizers divide a picture into equal-sized patches and pass the resulting sequence to a transformer. Flat sky receives the same token budget as a detailed face. QuadTok, developed by Matthew Mao and collaborators, replaces that fixed grid with a quadtree that adds tokens only where finer reconstruction improves perceptual quality.
| Benchmark | Result | Context |
|---|---|---|
| ImageNet reconstruction | About 10% fewer tokens | Roughly 230 tokens on average versus a fixed 256-token grid, at comparable reconstruction fidelity |
| COCO reconstruction | About 9% fewer tokens | Zero-shot evaluation using the ImageNet-trained tokenizer, without COCO fine-tuning |
| ImageNet generation | 2.08 gFID | A 947M-parameter GPT-style generator; lower gFID indicates closer agreement with the real-image distribution |
The Apache-2.0 release on GitHub contains tokenizer code and weights. The 947M-parameter generator checkpoint remains unpublished, leaving the reported generation result unavailable through official weights.
Detail gets the budget
Vector-quantized tokenizers used by models such as VAR, LlamaGen, and Parti convert image regions into discrete codebook entries that a transformer can process like vocabulary tokens. A common configuration maps a 256 × 256 image to a 16 × 16 grid, producing 256 tokens. During training, standard transformer attention grows with the square of sequence length; during autoregressive sampling, every additional token adds another sequential decoding step.
Compressed tokenizers such as TiTok and One-D-Piece produce shorter one-dimensional sequences. Their token positions have a weaker correspondence to specific image regions, which complicates coordinate-based conditioning and editing. QuadTok retains an explicit two-dimensional hierarchy while allowing each image to use a different number of tokens.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.