Skip to content
AI Primer
release

Perplexity releases pplx-embed-v2-late models for shared text-image retrieval

Perplexity's open pplx-embed-v2-late models retrieve text and images in a shared multi-vector space. The 0.6B model can query indexes built by the 9B model, including document pages without OCR.

4 min read
Perplexity releases pplx-embed-v2-late models for shared text-image retrieval
Perplexity releases pplx-embed-v2-late models for shared text-image retrieval

TL;DR

  • Perplexity released 0.6B and 9B multimodal embedding weights on Hugging Face, according to its announcement.
  • Cross-model querying is the standout feature: a 0.6B encoder can search a 9B-built index and recover roughly half the text-retrieval quality gap without additional query-encoding cost, per the team's results.
  • Rendered PDF pages are searchable without OCR, preserving layout, tables, and figures as the team described.
  • Agentic evaluations reached 64.0% answer accuracy on BrowseComp+ and 92.4% on MADQA with the 9B retriever, as the launch post reported.

The 9B checkpoint was the starting point for Perplexity's contextual embedding model. Perplexity also halved the small encoder's text tower and documented a token-marker mismatch with PyLate.

Shared embedding space

Perplexity distilled both sizes from the same 18B teacher using LEAF-style representation matching, according to the team's training thread. Aligning each retained token's vector with the teacher makes the two checkpoints' outputs interoperable.

Perplexity's deployment examples describe four configurations:

  1. Maximum quality: 9B encodes documents and queries.
  2. Fully local: 0.6B encodes documents and queries.
  3. Small-query, large-index: 9B encodes the corpus offline; 0.6B encodes live queries.
  4. Local-cloud: 0.6B encodes local documents or queries on-device, with representations comparable to a cloud-hosted 9B index.

The asymmetric setup changes the document encoder while keeping the query encoder fixed. The additional encoding work happens at indexing time, rather than on each request.

128-dimensional MaxSim

Each input retains one 128-dimensional vector per token. Unlike a cross-encoder, the model encodes documents independently of each query, as Perplexity explains in its architecture writeup.

Qwen3.5 encoders

Both models adapt Qwen3.5 with bidirectional attention. Perplexity documents two fine-tuning schemes in the 9B model card:

  • 0.6B: fully fine-tuned.
  • 9B: final eight transformer layers fully fine-tuned; remaining transformer layers and the vision encoder adapted with LoRA.

PDF pages and natural images

PDFs, slides, scans, and charts enter the retrieval pipeline as rendered images, avoiding the OCR and text-extraction losses described in the team's modality explanation. The visual-document evaluation uses the public subset of ViDoRe v3, with page images rather than extracted Markdown.

The 9B model trails the vision-specific EVIE models. Broader image evaluations cover Wikipedia screenshots through MIRACL-Vision and production-derived image queries through PPLX-Q2I; Gemini Embedding 2 leads Perplexity's 9B model on both.

Text and web retrieval

Training used 186 million query-document pairs from 594 datasets across 46 languages, with benchmark-associated datasets excluded, according to the team's data description. The 72-task text evaluation weights six domain groups equally, rather than weighting every task equally.

Health breaks the overall ranking: Nemotron scores 83.9 nDCG@10 against 82.5 for Perplexity's 9B model.

Q2D-Web tests roughly 70,000 agent-reformulated production queries against 190 million web documents. Perplexity reports Recall@1000 for this first-stage retrieval task, rather than the nDCG@10 used elsewhere in its evaluation.

Its Combined relevance judgments add LLM labels beyond documents surfaced by Perplexity's own citations and production retrieval system. The company acknowledges that the separate Citation and Web judgment sets favor models that rank similarly to those systems.

BrowseComp+ and MADQA

BrowseComp+ uses a GPT-OSS-120B agent with the tested retriever and another LLM to judge answers against ground truth, according to Perplexity's agentic evaluation. The 0.6B configuration issues fewer searches than the 9B configuration, while both beat the evaluated external retrievers on answer accuracy.

MADQA uses Gemini 3.5 Flash to answer 500 human-authored questions across 800 PDFs containing more than 18,000 pages. It measures answer accuracy and page-level F1 against annotated evidence pages.

The 0.6B advantage over the plain Mixedbread retriever falls within error bars. The 9B result also falls within the confidence interval of Mixedbread Agentic Search, which uses a search sub-agent running several searches per call.

Index storage

Late interaction stores document-token vectors and computes many token-to-token similarities, so storage and candidate-scoring costs grow with document length. Perplexity's evaluation writeup says it skipped EVIE-8B and Nemotron ColEmbed on the 100,000-image PPLX-Q2I corpus because indexing at their default output widths, up to 4,096 dimensions, was impractical.

The team pointed to hierarchical pooling and quantization as ways to reduce the deployment burden.

Sentence Transformers

MIT-licensed weights are available as 0.6B and 9B checkpoints, as tomaarsen confirmed. Native Sentence Transformers exports need no custom Python code.

Perplexity documents three integration details in the model card:

  • Package versions: sentence-transformers>=6.0.0 and transformers>=5.4.0.
  • Input batching: text-only and image-only batches require separate encoding calls; mixed text-plus-image inputs are unsupported.
  • PyLate markers: these models expect Q/D markers first; PyLate inserts them in the second position.

Further reading

Discussion across the web

Where this story is being discussed, in original context.

On X· 6 threads
TL;DR4 posts
Shared embedding space2 posts
128-dimensional MaxSim1 post
PDF pages and natural images1 post
Text and web retrieval1 post
Sentence Transformers1 post
Share on X