Cohere's Embed 5 Lets Developers Index Smart but Query 2.4x Faster

Cohere ships a two-tier embedding family with a shared vector space, 128K context, and a new benchmark methodology tuned for enterprise retrieval quality.

·
·
Cohere's Embed 5 Lets Developers Index Smart but Query 2.4x Faster
  • Cohere released Embed 5 in two tiers: Pro ($0.12/M tokens) and Fast ($0.08/M tokens).
  • Both share one embedding space, so index with Pro and query with Fast without re-indexing.
  • Pro averages 85.8 on ViDoRe V3, beating Voyage 4 Large, Gemini Embedding 2, and OpenAI text-embedding-3-large.
  • Supports 100+ languages, 128K context, multimodal inputs, Matryoshka dims (256-2048), and int8/binary outputs.
  • First model family optimized against RCP-nDCG@10, Cohere's new rubric-based retrieval metric.
  • Available via Cohere API, Model Vault, Microsoft Foundry, and Amazon SageMaker today.

Cohere Embed 5 pairs high-quality indexing with faster queries

Cohere has released Embed 5, an embedding-model family with two tiers that produce compatible vectors. Embed 5 Pro targets retrieval quality during offline corpus indexing, while Embed 5 Fast prioritizes latency and throughput for interactive search and agent loops. Because both models share an embedding space, developers can index documents with Pro and query them with Fast without rebuilding the vector index.

Embedding models convert text, images, and other content into numerical vectors whose proximity represents semantic similarity. Retrieval systems compare those vectors to find relevant documents, so Embed 5’s shared space separates the model used for ingestion from the model serving user requests. Cohere documents the models in its release notes.

One vector space, two operating points

Capability Embed 5 Pro Embed 5 Fast
Primary use High-quality offline indexing Low-latency, high-throughput queries
API list price $0.12 per million tokens $0.08 per million tokens
Reported throughput About 160 documents per second About 377 documents per second
Input formats Text, images, and combined text-and-image inputs such as PDF pages
Languages More than 100
Context window 128,000 tokens
Output dimensions 256, 512, 768, 1,024, 1,536, and 2,048
Output types Float, int8, and binary

Both variants support Matryoshka embeddings, which allow applications to store truncated vectors when lower storage use matters more than maximum retrieval quality. Cross-model retrieval requires the indexing and query paths to use the same output dimension and a representation supported by the vector database.

Embed 5 is available through Cohere’s Embed API, Microsoft Foundry, Amazon SageMaker, and Model Vault for single-tenant deployments. The listed per-token prices apply to the API; hosted and single-tenant deployment costs may follow platform-specific terms.

Vendor benchmarks favor Pro

Cohere reports that Embed 5 Pro averages 85.8 on ViDoRe V3, a benchmark built around visually complex enterprise documents. The company lists Embed 5 Fast at 84.7, Voyage 4 Large at 83.7, Gemini Embedding 2 at 83.2, and OpenAI’s text-embedding-3-large at 75.5.

Model ViDoRe V3 score
Embed 5 Pro 85.8
Embed 5 Fast 84.7
Voyage 4 Large 83.7
Gemini Embedding 2 83.2
OpenAI text-embedding-3-large 75.5

Finance-focused evaluations show a wider reported lead. Pro scores 80.1 on FinanceBench, 90.0 on FinQA, and 85.0 on ViDoRe V3 Finance, with Fast ranking second on each benchmark. Pro’s FinanceBench score exceeds OpenAI’s text-embedding-3-large result by 21.4 points. Compared with Embed 4, Cohere reports its largest multilingual gains in Farsi at 13 points, Telugu at 12 points, and Hindi at 12 points.

These results come from Cohere’s evaluation and should be read alongside workload-specific tests. Retrieval quality varies with chunking, document structure, language mix, query style, reranking, and the vector database’s similarity configuration.

Cross-model retrieval is the bet

Cohere evaluated every document-model and query-model pairing across 40 development datasets. Cross-model combinations averaged 1.6% and 2.7% lower performance than their corresponding same-model baselines, according to the company.

Those results support a two-stage deployment pattern: embed the corpus once with Pro, then encode each live query with Fast. Corpus ingestion receives Pro’s higher retrieval quality, while recurring requests use Fast’s lower price and higher throughput. Cohere measured Fast at an average of 2.4 times Pro’s document throughput across typical context lengths, or roughly 377 versus 160 documents per second.

Actual latency will depend on input length, batching, region, network overhead, and deployment hardware. Teams should benchmark end-to-end query time rather than infer request latency directly from Cohere’s document-throughput figures.

RCP-nDCG tackles missing labels

Cohere optimized Embed 5 using RCP-nDCG@10, an evaluation method designed to address incomplete relevance labels. Conventional normalized discounted cumulative gain, or nDCG, compares ranked results with a fixed answer key. A relevant document omitted from that key receives no credit when a retrieval system finds it.

In Cohere’s study with 46 human annotators, existing benchmark labels achieved an area under the curve, or AUC, of 0.65 when predicting human relevance judgments. A calibrated large-language-model judge reached 0.91. Human reviewers considered 28% of the documents marked irrelevant by the benchmark useful.

RCP-nDCG@10 applies a shared rubric to every retrieved document through a calibrated AI judge. Across pairwise system comparisons, the metric selected the system preferred by human reviewers 77% of the time, compared with 52% for conventional nDCG. Crediting relevant documents absent from the answer key accounted for a 19-point improvement.

Cohere has published the code and data in a GitHub repository. Results still depend on the judge model, calibration data, and rubric, and models optimized for RCP-nDCG may rank differently on legacy nDCG leaderboards.

Vectors can shrink to 32 bytes

Vector storage can exceed embedding-inference costs for large corpora. Embed 5 combines selectable dimensions with lower-precision output types, giving teams several storage profiles:

  • 2,048-dimensional float32: 8 KB per vector
  • 1,024-dimensional int8: 1 KB per vector
  • 256-dimensional binary: 32 bytes per vector

For 100 million chunks, those configurations require approximately 819 GB, 102 GB, and 3.2 GB of raw vector storage, respectively. The smallest option is 256 times smaller than a 2,048-dimensional float32 vector. Index metadata, graph structures, replicas, and database overhead add to those raw totals.

Cohere recommends 1,024-dimensional int8 vectors as a balance between size and retrieval quality, reporting near-float performance for both models. Database support varies, so applications need an index capable of storing and comparing the selected representation.

Index with Pro, query with Fast

The Python SDK exposes both models through the same embedding method. The document and query calls below use matching dimensions and float outputs while assigning the appropriate retrieval input type:

makefile
import os

import cohere

co = cohere.ClientV2(api_key=os.environ["CO_API_KEY"])

documents = co.embed(
    model="embed-v5.0-pro",
    input_type="search_document",
    texts=["Net interest margin narrowed 12 bps to 2.61%."],
    output_dimension=1024,
    embedding_types=["float"],
)

query = co.embed(
    model="embed-v5.0-fast",
    input_type="search_query",
    texts=["How did net interest margin change?"],
    output_dimension=1024,
    embedding_types=["float"],
)

document_vector = documents.embeddings.float_[0]
query_vector = query.embeddings.float_[0]

The document vector belongs in the corpus index, while the query vector is compared with stored vectors at request time. Existing retrieval pipelines must preserve the same dimension, output representation, and similarity configuration across both paths.

RAG systems working with parsed PDFs, financial filings, images, or multilingual corpora can use Pro for infrequent indexing jobs and Fast for repeated searches. The shared space removes the usual re-embedding step when switching between these two models, reducing migration time and avoiding a second copy of the corpus index.

Trending
  • No trending articles

Comments

avatar

Next Reads