Cohere's RCP-nDCG@10 Catches What Retrieval Benchmarks Miss

Cohere introduces Rubric-Calibrated Preferences nDCG@10, a new retrieval metric that uses a calibrated AI judge to score every result, not just those in the answer key.

·
·
·
Cohere's RCP-nDCG@10 Catches What Retrieval Benchmarks Miss
Read5 min
TypeNews
  • Cohere introduced RCP-nDCG@10, a retrieval metric using a calibrated AI judge instead of fixed answer keys.
  • Embed 5 is the first model family optimized and evaluated with this new methodology.
  • In human studies, RCP-nDCG agreed with reviewers 77% of the time versus 52% for standard nDCG.
  • Benchmark qrels label 29% of documents relevant where humans say 60%, exposing a huge coverage gap.
  • MTEB maintainer Kenneth Enevoldsen plans to integrate RCP-nDCG into the Massive Text Embedding Benchmark.
  • Code and data are open-sourced on GitHub for teams to run on their own corpus.

Cohere’s New Metric Targets Stale Retrieval Benchmarks

Top retrieval systems increasingly cluster within narrow leaderboard margins, making genuine improvements difficult to separate from better alignment with a benchmark’s labels. Cohere has proposed a new evaluation method, RCP-nDCG@10, and used it to evaluate and optimize its Embed 5 models.

RCP-nDCG@10, short for Rubric-Calibrated Preferences nDCG@10, grades each query-document pair with a calibrated LLM judge. The method combines a fixed relevance rubric with comparisons among retrieved documents. Cohere has released the accompanying code and data.

Static qrels leave documents behind

Normalized discounted cumulative gain, or nDCG, rewards systems for placing highly relevant documents near the top of a result list. Its relevance grades usually come from qrels, a fixed set of human judgments created when assessors review documents pooled from earlier retrieval systems.

Documents outside those pools may remain unjudged. In common evaluation setups, an unjudged document contributes no gain, even when it answers the query. Stronger retrievers can therefore lose credit for finding useful material that earlier systems missed.

Cohere measured the gap in a study involving 46 annotators:

  • Human reviewers found 28% of documents labeled irrelevant by the benchmark useful. Across the evaluated documents, benchmarks marked 29% as relevant, compared with 60% from human reviewers.
  • Existing qrels achieved an AUC of 0.65 when predicting human relevance judgments, compared with 0.91 for the calibrated LLM score. An AUC of 0.5 represents random discrimination between relevant and irrelevant documents.
  • Two reviewers independently grading the same document reached exact agreement on the zero-to-four relevance scale 42% of the time.

Missing judgments and reviewer disagreement impose an evaluation ceiling: once model quality exceeds the resolution of the labels, small leaderboard differences reveal little about which system produces better results.

A rubric and preferences share the work

RCP-nDCG@10 retains nDCG’s emphasis on ranking quality within the top 10 results. It replaces static relevance grades with two signals generated for each evaluated result list:

  1. Rubric judgments: The LLM answers five consistent yes-or-no questions for every query-document pair, producing an absolute relevance signal.
  2. Document preferences: The judge reviews documents in groups and estimates how likely each result is to outrank the others.
  3. Calibration: The method combines the rubric and preference signals into relevance grades used by nDCG@10.

The rubric anchors scores to a shared scale across queries, while the preference signal resolves the relative order of similar documents. This calibration limits the query-to-query drift that can occur when an LLM assigns a single relevance grade without a common reference.

Human reviewers favor the new score

Cohere tested the metric in 289 blind, head-to-head contests across NanoBEIR, BRIGHT, and ViDoRe v3. Three annotators reviewed each contest, and the majority preference determined the human-selected winner.

  • RCP-nDCG@10 selected the system preferred by reviewers in 77% of contests, compared with 52% for conventional nDCG.
  • When the metrics selected different winners, reviewers agreed with RCP-nDCG@10 in 70% of cases.
  • Among comparisons with the largest reported score margins, RCP-nDCG@10 agreed with reviewers 97% of the time. Conventional nDCG reached 53% at its largest margins.
  • Credit for relevant documents omitted from the qrels explained 19 percentage points of the 25-point improvement. Graded relevance accounted for the remaining six points.

MTEB maintainer Kenneth Enevoldsen said the team plans to add RCP-nDCG@10 to the MTEB leaderboard, citing stronger evaluation signals from existing datasets and lower runtime costs.

Embed 5 needs two scorecards

Embed 5 used RCP-nDCG@10 as an optimization target, so its reported results require attention to the retrieval stage and metric used. Cohere’s evaluation suite combines reranking, first-stage retrieval, multimodal retrieval, and cross-model tests.

Evaluation What it measures Reported metric
Reranking Ordering documents within a fixed candidate set RCP-nDCG@10
First-stage retrieval Finding candidates across the full corpus Standard nDCG and Recall
Fused text-image, page-image, and cross-model tests Ranking across multimodal or model-specific inputs Standard nDCG@10

A model’s position can change across these evaluations because each test measures a different part of the retrieval pipeline. Scores from separate rows in the table are therefore unsuitable for direct comparison.

Cohere reports a score of 83.4 for Embed 5 on its parsed-PDF evaluation, compared with 80.8 for Gemini Embedding 2 and 83.6 for Voyage 4 Large. Those figures belong to that specific evaluation setup.

Place it after candidate retrieval

Cohere’s GitHub repository provides the code and data needed to run RCP-nDCG@10 on another corpus. Suitable applications include:

  1. Evaluating proprietary RAG systems: Teams can score query-document pairs when annotated qrels are unavailable or expensive to maintain.
  2. Comparing embedding models: Evaluations can use the organization’s own queries, documents, and relevance criteria.
  3. Tracking retrieval regressions: Teams can compare model, chunking, and reranking changes as a corpus evolves.

RCP-nDCG@10 measures ranking quality inside a fixed candidate pool. Recall and first-stage nDCG remain necessary for checking whether candidate generation retrieved the relevant documents before reranking began.

Judge selection also becomes part of the evaluation configuration. Reproducible deployments should pin the judge model, prompt, rubric, and sampling settings; record evaluation costs; and verify a sample of judgments with human reviewers. Sensitive corpora require a judge deployment that meets the organization’s data-handling requirements.

Cohere’s search modeling team will discuss the methodology and its use in future releases during an X Live webinar, with registration available on the event page.

Trending
  • No trending articles

Comments

avatar

Next Reads