Hugging Face's ELECTRA Reranker Adds ONNX and OpenVINO for Faster CPU Search

A compact ELECTRA-based cross-encoder trained on MS MARCO for reranking search results, now shipping with ONNX, OpenVINO, and Safetensors weights.

·
·
Hugging Face's ELECTRA Reranker Adds ONNX and OpenVINO for Faster CPU SearchPRO
Read2 min
TypeModel
  • cross-encoder/ms-marco-electra-base refreshed on Hugging Face with new weight formats.
  • 0.1B-param reranker built on google/electra-base-discriminator, Apache-2.0 licensed, 908k downloads.
  • Now ships Safetensors, ONNX, and OpenVINO variants alongside PyTorch weights.
  • Scores 71.99 NDCG@10 on TREC DL 19 and 36.41 MRR@10 on MS MARCO Dev.
  • Runs at 340 docs/sec on a V100, slower than newer MiniLM-v2 checkpoints.
  • Best for second-stage passage reranking in English RAG and search pipelines.

ELECTRA reranker adds production-ready formats

The Hugging Face checkpoint for cross-encoder/ms-marco-electra-base now includes Safetensors, ONNX, and OpenVINO artifacts alongside its PyTorch weights. The underlying model remains the same; the expanded packaging gives developers more options for CPU inference, hardware acceleration, and safer weight loading.

Built from the ELECTRA base discriminator, the roughly 110 million-parameter model was fine-tuned for the English-language MS MARCO passage-ranking task. It uses the Apache 2.0 license and serves as a second-stage reranker for search and retrieval-augmented generation systems.

Why reranking cleans up top-k

First-stage retrievers such as BM25 and embedding indexes search large collections quickly, but their highest-ranked results often include weak matches. A cross-encoder rescoring step improves that ordering before the selected passages reach a search interface or language model.

Bi-encoder retrieval embeds queries and passages separately, then compares their vectors. A cross-encoder feeds each query-passage pair through one transformer, allowing every query token to attend to every passage token before the model produces a relevance score. That joint processing improves ranking precision while increasing computation because the model must run once for every candidate.

  1. Retrieve a candidate set, commonly 20 to 100 passages.
  2. Pair each passage with the original query.
  3. Score the pairs in batches with the cross-encoder.
  4. Sort by score and keep the highest-ranked passages.

Run it in a few lines

The Sentence Transformers wrapper handles pairwise tokenization, batching, device placement, and score conversion:

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads