LAION's Open CLIP Model Hits 70% ImageNet Accuracy With 150M Parameters
LAION's open ViT-B/16 CLIP, trained on 2 billion image-text pairs, hits 70.2% zero-shot on ImageNet and powers nearly 100 spaces.
- LAION's open ViT-B/16 CLIP hits 70.2% zero-shot top-1 on ImageNet-1k, trained on 2B LAION-2B pairs.
- Trained with 34B samples seen at 88K batch size on the JUWELS Booster supercomputer.
- Loadable in two lines via OpenCLIP, shipped in safetensors under MIT license.
- Powers ~100 Hugging Face Spaces including PuLID-FLUX and other generative conditioning pipelines.
- DataComp.XL variant of same architecture reaches 73.5% with better data curation.
- Explicitly research-only: not for deployment, surveillance, or non-English use.
Why LAION’s ViT-B/16 CLIP checkpoint stays busy
LAION’s ViT-B/16 checkpoint continues to record hundreds of thousands of monthly downloads on Hugging Face and appears in many downstream Spaces. The model is an independent training run of OpenAI’s CLIP architecture, built from scratch on roughly 2 billion image-text pairs with the OpenCLIP framework. Its combination of open weights, established tooling, moderate hardware requirements, and known benchmark performance keeps it useful as a vision-language baseline.
The checkpoint, decoded
| Architecture | ViT-B/16 image encoder paired with a Transformer text encoder |
|---|---|
| Image input | 224 × 224 pixels divided into 16 × 16 patches |
| Projection size | 512-dimensional image and text embeddings |
| Text context | Up to 77 tokens with the checkpoint’s tokenizer |
| Training data | LAION-2B English, a subset of LAION-5B derived from image URLs and nearby text collected from public webpages |
| Training exposure | 34 billion samples processed with a global batch size of about 88,000 |
| Training system | JUWELS Booster supercomputer |
| Reported benchmark | 70.2% zero-shot top-1 accuracy on ImageNet-1K |
| Size | About 150 million parameters and roughly 600 MB for FP32 weights |
| License | MIT, alongside deployment guidance and risk warnings in the model card |
The suffix s34B-b88K records the training run’s sample exposure and batch size. The 34 billion figure counts examples processed during training, including repeated passes, rather than unique image-text pairs. CLIP’s contrastive objective raises the similarity of matched images and captions while separating mismatched examples from the same batch.
Where B/16 lands
LAION released several OpenCLIP sizes with different accuracy and compute profiles. At 224-pixel resolution, B/16 creates 196 image-patch tokens, compared with 49 for B/32. The finer grid improves visual detail while increasing attention cost. Larger L/14 and H/14 encoders add capacity and require substantially more memory and compute.
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.