Perplexity's pplx-embed-v2 Beats Voyage at 8x Smaller Vector Size
Perplexity's new 9B contextual embedding model distills relevance from a compression teacher, beating Voyage on turbopuffer's private context-bench with 8x less storage.
- Perplexity released pplx-embed-v2-context-9b-preview, a 9B contextual embedding model with 1024/2048 dim and native int8 support.
- Training distills chunk relevance from a query-aware context compression model instead of using single gold-chunk labels.
- Leads context-bench on Answer and Evidence recall at every cutoff, beating voyage-context-4 by 14.4 points at K=10.
- Matches voyage-context-4 quality using 1 KB per vector vs 8 KB, an 8x storage reduction on chunk retrieval.
- Highest average nDCG@10 on public ConTEB benchmark across contextual embedding models tested.
- Context-bench, held privately by turbopuffer, has 2,099 queries across 38,894 long documents in 21 domains.
Perplexity previews a 9B contextual embedding model with 1 KB vectors
Perplexity has released preview weights for pplx-embed-v2-context-9b-preview, a contextual embedding model that encodes each chunk with its full document in view. Perplexity reports the highest average score on ConTEB and leading results on context-bench, a private benchmark created with turbopuffer.
The model produces 1024-dimensional and int8 embeddings. A raw 1,024-dimensional int8 vector occupies 1,024 bytes, excluding index structures and metadata.
Why isolated chunks lose meaning
Retrieval-augmented systems divide long documents into independently searchable chunks. A chunk can lose the heading, entity definition, table header, speaker name, or date that gives its text meaning. Documents with repeated language, such as leases or regulatory filings, expose the problem because several chunks may look identical without their surrounding context.
Contextual models use late chunking to preserve that information. The model processes the document in one forward pass, then pools chunk vectors from contextualized token representations. Each resulting vector reflects information from elsewhere in the document while remaining available for independent indexing and retrieval.
Traditional training often depends on gold-chunk annotations generated by an LLM. The annotator selects one chunk as the answer, and a contrastive loss treats the document’s other chunks as negatives. That approach discards useful supporting passages, costs more as the corpus grows, and binds supervision to the annotator’s chosen chunk boundaries.
Token scores replace fixed gold chunks
Perplexity trains the embedding model with a compression teacher that reads a query and document together. The teacher assigns every token a continuous relevance score, which can be aggregated across any chunk boundary. Randomized chunking lets the same token-level signal supervise multiple document splits.
Two losses shape the student
- Document-level InfoNCE: The system defines document similarity as the highest cosine similarity between the query and any chunk in that document. The loss moves the matching document above other documents in the batch. This resembles ColBERT’s MaxSim operation at chunk level.
- Chunk-level KL distillation: The teacher’s token scores become a target probability distribution over chunks in the positive document. The student learns to match that distribution with its softmax over all chunks in the batch.
The query-aware teacher supplies training targets only. Indexed document vectors can still be computed before search, without running the compression model for each user query.
Perplexity initializes the student from an in-house 9B-parameter ColBERT retrieval model and adds a linear projection that produces 2,048-dimensional embeddings. Matryoshka training supports nested 1,024- and 2,048-dimensional representations, while quantization-aware training prepares the vectors for native int8 output.
The training mixture contains roughly 430 public and internal query-document datasets across more than 50 languages. Perplexity says it excluded ConTEB and context-bench data during development.
A benchmark for ambiguous matches
Context-bench tests retrieval cases in which local wording cannot reliably identify the correct source. Context-bench consists of 2,099 queries over 38,894 long documents in 21 domains. Primary target documents have a median length of about 6,100 tokens.
| Measure | Count |
|---|---|
| Queries | 2,099 |
| Documents | 38,894 |
| Sentence chunks | 2,458,072 |
| Domains | 21 |
| Contextual capabilities | 12 |
The twelve capabilities include entity identity, table structure, pronoun and reference resolution, list position, speaker attribution, changes over time, causal relationships, and cross-language meaning. A property-management query, for example, may need to distinguish thousands of leases containing the same sentence pattern, such as “Monthly rent is $X,XXX.” Distant details about the address, tenant, and dates identify the correct lease.
Three views of retrieval quality
- Document@K: Measures whether the correct document ranks among the top K results when near-identical documents share vocabulary.
- Answer@K: Measures whether an answer-containing chunk appears among the top K chunks across the corpus.
- Evidence Recall@K: Removes the leading answer chunk, then measures how much supporting evidence appears among the top K chunks from the correct document.
The benchmark corpus remains private to reduce training contamination. That design limits independent inspection, so published context-bench scores depend on turbopuffer’s controlled evaluation process.
Where the preview leads
| Metric | Score | Comparison |
|---|---|---|
| Answer Recall@10 | 45.5% | 14.4 percentage points above voyage-context-4 |
| Evidence Recall@10 | 40.6% | 5.0 percentage points above voyage-context-4 |
| All-Evidence@10 | 31.1% | Highest reported result |
| Document@1 | 15.2% | Highest reported result |
| Document@10 | 61.6% | Highest reported result |
The preview leads the reported Answer@K and Evidence Recall@K results at every evaluated cutoff. Perplexity’s earlier pplx-context-v1-4B retains higher Document@3 and Document@5 scores, showing that the new model does not dominate every document-ranking measure.
On public ConTEB evaluations, the preview records the highest average nDCG@10, a ranking metric that rewards placing relevant results near the top. pplx-context-v1-4B scores higher on NarrativeQA, while Nemotron-3-8B leads on COVID-QA. Perplexity attributes the COVID-QA result to medical terms that allow stronger lexical matching with less dependence on document context.
One kilobyte per vector
| Configuration | Dimensions | Format | Bytes per vector |
|---|---|---|---|
| pplx-embed-v2 preview | 1,024 | int8 | 1,024 |
| voyage-context-4 comparison | 2,048 | float32 | 8,192 |
The compact configuration uses one-eighth of the raw vector storage of the cited voyage-context-4 configuration and slightly exceeds its reported average chunk-retrieval score. Actual index size will also include document metadata, identifiers, alignment, and approximate-nearest-neighbor data structures.
Reported chunk-size sensitivity is modest across the tested range. Average nDCG@10 declines from 81.0% with 64-token chunks to 79.9% with 512-token chunks, giving indexing pipelines room to choose splits based on document structure and serving costs.
Operational changes for retrieval pipelines
- Encode complete documents: Existing workers that embed each chunk independently must preserve document grouping and run the full document through the encoder before pooling chunk vectors.
- Re-embed updated documents: A change in one section can affect contextual representations elsewhere, so an update may require regenerating every chunk vector for that document.
- Check int8 support: Vector databases and approximate-nearest-neighbor indexes must support 1,024-dimensional int8 vectors to retain the advertised storage benefit. Conversion to wider numeric formats increases memory use.
- Benchmark local data: Contextual retrieval helps most when relevant passages depend on distant information. Workloads driven by distinctive local terms may see smaller gains.
- Measure indexing costs: Whole-document encoding changes batching, maximum-length handling, throughput, and GPU memory requirements compared with independent chunk embedding.
Self-hosted preview weights are available now, and Perplexity says API access will follow without giving a release date. Context-bench submissions require coordination with turbopuffer because the evaluation corpus is private.
Production evaluations still need to establish maximum input length, latency, throughput, hardware requirements, license terms, and total index overhead. The benchmark results and raw vector sizes provide retrieval and storage signals, while those deployment characteristics will determine the model’s fit in a live system.