Hugging Face's ELECTRA Reranker Adds ONNX and OpenVINO for Faster CPU Search
A compact ELECTRA-based cross-encoder trained on MS MARCO for reranking search results, now shipping with ONNX, OpenVINO, and Safetensors weights.
- cross-encoder/ms-marco-electra-base refreshed on Hugging Face with new weight formats.
- 0.1B-param reranker built on google/electra-base-discriminator, Apache-2.0 licensed, 908k downloads.
- Now ships Safetensors, ONNX, and OpenVINO variants alongside PyTorch weights.
- Scores 71.99 NDCG@10 on TREC DL 19 and 36.41 MRR@10 on MS MARCO Dev.
- Runs at 340 docs/sec on a V100, slower than newer MiniLM-v2 checkpoints.
- Best for second-stage passage reranking in English RAG and search pipelines.
ELECTRA reranker adds production-ready formats
The Hugging Face checkpoint for cross-encoder/ms-marco-electra-base now includes Safetensors, ONNX, and OpenVINO artifacts alongside its PyTorch weights. The underlying model remains the same; the expanded packaging gives developers more options for CPU inference, hardware acceleration, and safer weight loading.
Built from the ELECTRA base discriminator, the roughly 110 million-parameter model was fine-tuned for the English-language MS MARCO passage-ranking task. It uses the Apache 2.0 license and serves as a second-stage reranker for search and retrieval-augmented generation systems.
Why reranking cleans up top-k
First-stage retrievers such as BM25 and embedding indexes search large collections quickly, but their highest-ranked results often include weak matches. A cross-encoder rescoring step improves that ordering before the selected passages reach a search interface or language model.
Bi-encoder retrieval embeds queries and passages separately, then compares their vectors. A cross-encoder feeds each query-passage pair through one transformer, allowing every query token to attend to every passage token before the model produces a relevance score. That joint processing improves ranking precision while increasing computation because the model must run once for every candidate.
- Retrieve a candidate set, commonly 20 to 100 passages.
- Pair each passage with the original query.
- Score the pairs in batches with the cross-encoder.
- Sort by score and keep the highest-ranked passages.
Run it in a few lines
The Sentence Transformers wrapper handles pairwise tokenization, batching, device placement, and score conversion:
This story is for Pro members
You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.