Developers Are Rediscovering e5-small-v2's 33M-Parameter Search Power
A 33M parameter English embedding model is climbing Hugging Face trending charts, offering 384 dimensional vectors with surprisingly competitive MTEB scores.
- intfloat/e5-small-v2, a 33.4M parameter English embedding model, is trending on Hugging Face with 700K+ monthly downloads.
- Produces 384 dimensional vectors from 12 BERT layers, licensed MIT for commercial use.
- Trained via weakly supervised contrastive pre-training then supervised fine tuning, described in arXiv 2212.03533.
- Requires query: and passage: prefixes on inputs or performance degrades noticeably.
- Hits 91.3 accuracy on Amazon polarity and competitive scores across MTEB, rivaling much larger encoders.
- English only, 512 token cap, with 41 fine tunes and 16 quantized variants in its ecosystem.
Why e5-small-v2 is trending again
intfloat/e5-small-v2 has returned to Hugging Face’s trending list, with its model page showing more than 700,000 monthly downloads at publication. Hugging Face counts artifact requests, which can include automated builds and repeated pulls, so the figure signals activity rather than a tally of unique users.
Liang Wang and collaborators introduced the E5 family in the E5 paper. The renewed interest matters because this 33.4 million-parameter encoder gives developers a fast, inexpensive retrieval baseline while newer embedding systems consume far more memory or require paid APIs.
What 33 million parameters buy
| Property | Value |
|---|---|
| Architecture | 12-layer BERT-style bi-encoder |
| Parameters | 33.4 million |
| Embedding width | 384 dimensions |
| Maximum input | 512 tokens; longer inputs are truncated |
| Language | English |
| License | MIT |
| Deployment options | Sentence Transformers, Transformers, PyTorch, TensorFlow, ONNX, Safetensors, and OpenVINO |
A bi-encoder processes queries and documents separately, allowing an application to embed its document collection once and search it with fast vector comparisons. The 384-dimensional output also reduces storage and index memory compared with encoders that produce 768-dimensional vectors.
The prefixes are part of the model
Sentence Transformers provides the shortest path to a working implementation. Every input requires a task prefix followed by a space:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/e5-small-v2")
texts = [
"query: how much protein should a female eat",
"passage: the CDC's average requirement of protein for women ages 19 to 70 is 46 grams per day",
]
embeddings = model.encode(
texts,
normalize_embeddings=True,
)
scores = (embeddings[:1] @ embeddings[1:].T) * 100This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.