LAION's Open CLIP Model Hits 70% ImageNet Accuracy With 150M Parameters

LAION's open ViT-B/16 CLIP, trained on 2 billion image-text pairs, hits 70.2% zero-shot on ImageNet and powers nearly 100 spaces.

·
·
·
LAION's Open CLIP Model Hits 70% ImageNet Accuracy With 150M ParametersPRO
Read2 min
TypeModel
  • LAION's open ViT-B/16 CLIP hits 70.2% zero-shot top-1 on ImageNet-1k, trained on 2B LAION-2B pairs.
  • Trained with 34B samples seen at 88K batch size on the JUWELS Booster supercomputer.
  • Loadable in two lines via OpenCLIP, shipped in safetensors under MIT license.
  • Powers ~100 Hugging Face Spaces including PuLID-FLUX and other generative conditioning pipelines.
  • DataComp.XL variant of same architecture reaches 73.5% with better data curation.
  • Explicitly research-only: not for deployment, surveillance, or non-English use.

Why LAION’s ViT-B/16 CLIP checkpoint stays busy

LAION’s ViT-B/16 checkpoint continues to record hundreds of thousands of monthly downloads on Hugging Face and appears in many downstream Spaces. The model is an independent training run of OpenAI’s CLIP architecture, built from scratch on roughly 2 billion image-text pairs with the OpenCLIP framework. Its combination of open weights, established tooling, moderate hardware requirements, and known benchmark performance keeps it useful as a vision-language baseline.

The checkpoint, decoded

Architecture ViT-B/16 image encoder paired with a Transformer text encoder
Image input 224 × 224 pixels divided into 16 × 16 patches
Projection size 512-dimensional image and text embeddings
Text context Up to 77 tokens with the checkpoint’s tokenizer
Training data LAION-2B English, a subset of LAION-5B derived from image URLs and nearby text collected from public webpages
Training exposure 34 billion samples processed with a global batch size of about 88,000
Training system JUWELS Booster supercomputer
Reported benchmark 70.2% zero-shot top-1 accuracy on ImageNet-1K
Size About 150 million parameters and roughly 600 MB for FP32 weights
License MIT, alongside deployment guidance and risk warnings in the model card

The suffix s34B-b88K records the training run’s sample exposure and batch size. The 34 billion figure counts examples processed during training, including repeated passes, rather than unique image-text pairs. CLIP’s contrastive objective raises the similarity of matched images and captions while separating mismatched examples from the same batch.

Where B/16 lands

LAION released several OpenCLIP sizes with different accuracy and compute profiles. At 224-pixel resolution, B/16 creates 196 image-patch tokens, compared with 49 for B/32. The finer grid improves visual detail while increasing attention cost. Larger L/14 and H/14 encoders add capacity and require substantially more memory and compute.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads