Perplexity's pplx-embed-v2-late Searches 190M Docs Without OCR or Chunking
Perplexity's new late-interaction embedding family keeps 128-dim vectors per token, searches rendered pages without OCR, and shares one space across sizes.
- Perplexity released pplx-embed-v2-late, 0.6B and 9B late-interaction embedding models for text, images, and rendered pages.
- ColBERT-style architecture outputs 128-dim vectors per token, scored with MaxSim for finer query-document matching.
- Shared embedding space lets a 0.6B query encoder search a 9B-built index, lifting ViDoRe v3 from 62.3% to 63.5%.
- Both models embed PDFs and slides directly, skipping OCR while preserving tables, figures, and layout.
- 9B hits 74.8% Recall@1000 on Q2D-Web and 92.4% on MADQA, beating Mixedbread's retriever by 3.5pp.
- Agentic BrowseComp+ accuracy reaches 64.0% with GPT-OSS-120B, 4.9pp over the next ColBERT model with fewer searches.
Perplexity has released pplx-embed-v2-late, two late-interaction embedding models that place text, natural images, and rendered document pages in a shared embedding space. The 0.6B and 9B checkpoints are available in a Hugging Face collection, support sentence-transformers >= 6.0.0, and produce compatible embeddings, allowing queries encoded by the smaller model to search an index built by the larger one.
Each token keeps its own vector
Most dense retrieval systems compress each indexed item into one vector, then compare it with a query vector using an inexpensive inner product. That design makes billion-scale approximate nearest-neighbor search practical, but longer documents must squeeze more topics and details into the same representation. Chunking reduces the compression pressure while risking broken context across paragraphs, tables, figures, and page boundaries.
Late-interaction models encode queries and documents independently, then compare their token-level representations during scoring. The pplx-embed-v2-late models emit a 128-dimensional vector for every retained token and use a ColBERT-style MaxSim calculation:
- Encode the query and document as sequences of vectors.
- For each query vector, find the most similar document vector.
- Sum those maximum similarities into one retrieval score.
This scoring method lets separate query terms match different passages or visual regions without compressing the entire document into one vector.
PDF pages can bypass OCR
A conventional document pipeline parses PDFs, runs OCR or captioning, and indexes the extracted text. Those steps add processing cost and can lose table structure, figure labels, slide layouts, and spatial relationships. Pplx-embed-v2-late can instead encode a rendered page directly and compare its visual tokens with a text query in the same embedding space.
A PDF still needs to be rendered as an image before encoding, and retrieval quality remains sensitive to page resolution and scan quality. Bypassing text extraction removes OCR as a separate source of errors, while downstream answer generation still requires a vision-capable model or a later extraction step if the selected pages must be converted into text.
A shared teacher aligns both sizes
Perplexity first trains an 18B teacher with a contrastive objective that pulls matching queries and documents together while separating irrelevant pairs. It then distills that teacher independently into the 9B and 0.6B students using LEAF-style representation distillation.
The distillation process aligns each student’s output with the teacher’s vector for every retained token. Because both students learn the same representation targets, their vectors occupy compatible coordinates, enabling the 0.6B model to score queries against documents encoded by the 9B model.
The smaller checkpoint begins with Qwen3.5-0.8B and reduces its text tower from 24 layers to 12. That change cuts the tower from 498 million to 240 million non-embedding parameters. The finished model contains 594 million parameters, with about 240 million active for text encoding and 340 million for image encoding, so each modality invokes only part of the model.
Deployment can mix model sizes
| Configuration | Document encoder | Query encoder | Best fit |
|---|---|---|---|
| Quality-first | 9B | 9B | Highest reported retrieval scores when query latency and memory allow it |
| Local | 0.6B | 0.6B | Smaller deployments and fully local inference |
| Asymmetric | 9B | 0.6B | Offline high-quality indexing with lower live-query cost |
| Device-cloud split | 9B in the cloud | 0.6B on device | Local query encoding before transmitting embeddings for search |
Local encoding can keep raw inputs on the device, although transmitted embeddings may still reveal information and require the same security review as other derived data.
Production use requires a multi-vector retrieval backend. Sentence Transformers covers model loading and encoding, but the index must preserve token vectors and implement MaxSim scoring; a standard one-vector approximate nearest-neighbor schema cannot consume these outputs unchanged.
Reported gains span text and vision
Perplexity evaluated the models on domain-specific text, web search, and visual document retrieval. The company says it excluded datasets associated with every reported benchmark from the training mix. The results remain vendor-reported, although the downloadable checkpoints allow independent testing.
| Benchmark | Scale and metric | 9B | 0.6B | Comparison |
|---|---|---|---|---|
| Domain suite | 72 tasks across finance, law, health, conversation, technology, and multilingual retrieval | 81.3 | 78.0 | 9B leads the next model by 1.6 points; 0.6B trails Gemini Embedding 2 by 0.3 |
| Q2D-Web | About 70,000 queries over 190 million documents, measured with Recall@1000 | 74.8% | 73.6% | Previous best: 69.3% |
| ViDoRe v3 | Average visual document retrieval score | 65.2 | 62.3 | 9B trails Tencent EVIE; 0.6B leads topk-embed-v1-2b |
Q2D-Web uses production-derived queries, and Recall@1000 measures how much relevant material appears among the first 1,000 retrieved results. Both Perplexity models lead the comparison set on its combined judgments.
The asymmetric configuration preserves the 0.6B query-time workload while improving document representations. Encoding documents with 9B and queries with 0.6B raises the domain-suite score by 1.6 points over using 0.6B on both sides, and lifts ViDoRe v3 from 62.3 to 63.5.
Agents retrieve with fewer searches
BrowseComp+ measures whether an LLM can gather enough retrieved evidence to answer complex questions. Perplexity reports 64.0% accuracy for the 9B model, 4.9 percentage points above the next ColBERT model and 8.7 points above the next dense model. It also issues fewer searches than every evaluated model except the 0.6B checkpoint.
MADQA contains 500 questions over 800 real-world PDFs totaling more than 18,000 pages. The 9B model reaches 92.4% accuracy, 3.5 points above the Mixedbread retriever, while the 0.6B model reaches 90.1%.
Higher recall enlarges the index
Late interaction increases storage and scoring work because every indexed item retains multiple vectors. Dense retrieval commonly stores one vector per item and performs one similarity calculation per candidate; MaxSim evaluates many query-to-document token pairs, with costs that grow alongside document length.
Perplexity limits each token vector to 128 dimensions, compared with 2,048 or 4,096 dimensions for the cited vision-only competitors. Actual index size still depends on retained token counts, numerical precision, and compression, so model parameter counts alone do not predict storage or serving costs.
Document encoding with the 9B checkpoint can run offline, making the asymmetric setup practical when reindexing time is acceptable. Compatibility should be assumed only for the aligned pplx-embed-v2-late checkpoints, since unrelated embedding models do not automatically share the same vector space.
Where the design pays off
- PDFs, scans, and slide decks: Direct page encoding preserves visual structure and can remove OCR from the retrieval stage.
- Long or multi-topic documents: Token-level scoring retains details that a single vector may compress away.
- Latency-sensitive search: A 9B document index paired with 0.6B query encoding moves most computation into an offline job.
- Agentic retrieval: Better ranking can reduce repeated searches and the model calls attached to them.
- Storage-constrained systems: A dense model may remain simpler and less expensive when multi-vector gains do not justify a larger index.
At launch, both checkpoints are available for local use through Hugging Face and Sentence Transformers. Perplexity says API access will follow alongside dense and contextual embedding variants in the coming months.