Perplexity's pplx-embed-v2 Beats Voyage at 8x Smaller Vector Size

Perplexity's new 9B contextual embedding model distills relevance from a compression teacher, beating Voyage on turbopuffer's private context-bench with 8x less storage.

·
·
Perplexity's pplx-embed-v2 Beats Voyage at 8x Smaller Vector Size
Read6 min
TypeNews
SubtopicEmbeddings · Rag
  • Perplexity released pplx-embed-v2-context-9b-preview, a 9B contextual embedding model with 1024/2048 dim and native int8 support.
  • Training distills chunk relevance from a query-aware context compression model instead of using single gold-chunk labels.
  • Leads context-bench on Answer and Evidence recall at every cutoff, beating voyage-context-4 by 14.4 points at K=10.
  • Matches voyage-context-4 quality using 1 KB per vector vs 8 KB, an 8x storage reduction on chunk retrieval.
  • Highest average nDCG@10 on public ConTEB benchmark across contextual embedding models tested.
  • Context-bench, held privately by turbopuffer, has 2,099 queries across 38,894 long documents in 21 domains.

Perplexity previews a 9B contextual embedding model with 1 KB vectors

Perplexity has released preview weights for pplx-embed-v2-context-9b-preview, a contextual embedding model that encodes each chunk with its full document in view. Perplexity reports the highest average score on ConTEB and leading results on context-bench, a private benchmark created with turbopuffer.

The model produces 1024-dimensional and int8 embeddings. A raw 1,024-dimensional int8 vector occupies 1,024 bytes, excluding index structures and metadata.

Why isolated chunks lose meaning

Retrieval-augmented systems divide long documents into independently searchable chunks. A chunk can lose the heading, entity definition, table header, speaker name, or date that gives its text meaning. Documents with repeated language, such as leases or regulatory filings, expose the problem because several chunks may look identical without their surrounding context.

Contextual models use late chunking to preserve that information. The model processes the document in one forward pass, then pools chunk vectors from contextualized token representations. Each resulting vector reflects information from elsewhere in the document while remaining available for independent indexing and retrieval.

Traditional training often depends on gold-chunk annotations generated by an LLM. The annotator selects one chunk as the answer, and a contrastive loss treats the document’s other chunks as negatives. That approach discards useful supporting passages, costs more as the corpus grows, and binds supervision to the annotator’s chosen chunk boundaries.

Token scores replace fixed gold chunks

Perplexity trains the embedding model with a compression teacher that reads a query and document together. The teacher assigns every token a continuous relevance score, which can be aggregated across any chunk boundary. Randomized chunking lets the same token-level signal supervise multiple document splits.

Training pipeline connecting a student embedding model to a query-aware compression teacher
The compression model supplies token-level relevance targets for the student embedder.

Two losses shape the student

  • Document-level InfoNCE: The system defines document similarity as the highest cosine similarity between the query and any chunk in that document. The loss moves the matching document above other documents in the batch. This resembles ColBERT’s MaxSim operation at chunk level.
  • Chunk-level KL distillation: The teacher’s token scores become a target probability distribution over chunks in the positive document. The student learns to match that distribution with its softmax over all chunks in the batch.

The query-aware teacher supplies training targets only. Indexed document vectors can still be computed before search, without running the compression model for each user query.

Perplexity initializes the student from an in-house 9B-parameter ColBERT retrieval model and adds a linear projection that produces 2,048-dimensional embeddings. Matryoshka training supports nested 1,024- and 2,048-dimensional representations, while quantization-aware training prepares the vectors for native int8 output.

The training mixture contains roughly 430 public and internal query-document datasets across more than 50 languages. Perplexity says it excluded ConTEB and context-bench data during development.

A benchmark for ambiguous matches

Context-bench tests retrieval cases in which local wording cannot reliably identify the correct source. Context-bench consists of 2,099 queries over 38,894 long documents in 21 domains. Primary target documents have a median length of about 6,100 tokens.

Context-bench scope
Measure Count
Queries 2,099
Documents 38,894
Sentence chunks 2,458,072
Domains 21
Contextual capabilities 12

The twelve capabilities include entity identity, table structure, pronoun and reference resolution, list position, speaker attribution, changes over time, causal relationships, and cross-language meaning. A property-management query, for example, may need to distinguish thousands of leases containing the same sentence pattern, such as “Monthly rent is $X,XXX.” Distant details about the address, tenant, and dates identify the correct lease.

Three views of retrieval quality

  • Document@K: Measures whether the correct document ranks among the top K results when near-identical documents share vocabulary.
  • Answer@K: Measures whether an answer-containing chunk appears among the top K chunks across the corpus.
  • Evidence Recall@K: Removes the leading answer chunk, then measures how much supporting evidence appears among the top K chunks from the correct document.

The benchmark corpus remains private to reduce training contamination. That design limits independent inspection, so published context-bench scores depend on turbopuffer’s controlled evaluation process.

Where the preview leads

Context-bench comparison of answer recall, document retrieval, and evidence recall
Perplexity’s reported context-bench results across answer, document, and evidence metrics.
Reported context-bench results
Metric Score Comparison
Answer Recall@10 45.5% 14.4 percentage points above voyage-context-4
Evidence Recall@10 40.6% 5.0 percentage points above voyage-context-4
All-Evidence@10 31.1% Highest reported result
Document@1 15.2% Highest reported result
Document@10 61.6% Highest reported result

The preview leads the reported Answer@K and Evidence Recall@K results at every evaluated cutoff. Perplexity’s earlier pplx-context-v1-4B retains higher Document@3 and Document@5 scores, showing that the new model does not dominate every document-ranking measure.

On public ConTEB evaluations, the preview records the highest average nDCG@10, a ranking metric that rewards placing relevant results near the top. pplx-context-v1-4B scores higher on NarrativeQA, while Nemotron-3-8B leads on COVID-QA. Perplexity attributes the COVID-QA result to medical terms that allow stronger lexical matching with less dependence on document context.

One kilobyte per vector

Retrieval quality plotted against raw storage per embedding vector
The 1,024-dimensional int8 configuration reduces raw vector storage while preserving the reported retrieval average.
Raw vector payload comparison
Configuration Dimensions Format Bytes per vector
pplx-embed-v2 preview 1,024 int8 1,024
voyage-context-4 comparison 2,048 float32 8,192

The compact configuration uses one-eighth of the raw vector storage of the cited voyage-context-4 configuration and slightly exceeds its reported average chunk-retrieval score. Actual index size will also include document metadata, identifiers, alignment, and approximate-nearest-neighbor data structures.

Reported chunk-size sensitivity is modest across the tested range. Average nDCG@10 declines from 81.0% with 64-token chunks to 79.9% with 512-token chunks, giving indexing pipelines room to choose splits based on document structure and serving costs.

Operational changes for retrieval pipelines

  • Encode complete documents: Existing workers that embed each chunk independently must preserve document grouping and run the full document through the encoder before pooling chunk vectors.
  • Re-embed updated documents: A change in one section can affect contextual representations elsewhere, so an update may require regenerating every chunk vector for that document.
  • Check int8 support: Vector databases and approximate-nearest-neighbor indexes must support 1,024-dimensional int8 vectors to retain the advertised storage benefit. Conversion to wider numeric formats increases memory use.
  • Benchmark local data: Contextual retrieval helps most when relevant passages depend on distant information. Workloads driven by distinctive local terms may see smaller gains.
  • Measure indexing costs: Whole-document encoding changes batching, maximum-length handling, throughput, and GPU memory requirements compared with independent chunk embedding.

Self-hosted preview weights are available now, and Perplexity says API access will follow without giving a release date. Context-bench submissions require coordination with turbopuffer because the evaluation corpus is private.

Production evaluations still need to establish maximum input length, latency, throughput, hardware requirements, license terms, and total index overhead. The benchmark results and raw vector sizes provide retrieval and storage signals, while those deployment characteristics will determine the model’s fit in a live system.

Trending
  • No trending articles

Comments

avatar

Next Reads