Perplexity open-sources pplx-embed-v2-context-9b-preview
Perplexity released pplx-embed-v2-context-9b-preview, which encodes chunks using whole-document context. Perplexity reports leading results on ConTEB and Turbopuffer's context benchmark.

TL;DR
- Perplexity open-sourced a preview whose chunks are encoded with the whole document in view, according to the open-source announcement.
- The training target replaces a single gold chunk with teacher-derived relevance scores for answer and supporting chunks, according to the training summary.
- On turbopuffer's private context-bench, Perplexity reports 45.5% answer recall@10, 40.6% evidence recall@10, and 61.6% document recall@10, with the answer score 14.4 points above voyage-context-4 the benchmark report.
- The 1024-dimensional int8 format uses 1 KB per vector and still beats voyage-context-4's 2048-dimensional float32 format at 8 KB per vector, according to the storage comparison.
- The release is a self-hosted Hugging Face preview, while Perplexity says API availability is still in progress the API follow-up.
Query and document encoding use separate methods, and the model card warns that sending queries through the document path silently degrades retrieval. The technical post pairs the model with a benchmark whose documents and queries remain private, while the benchmark thread shows the storage and retrieval comparisons. The teacher runs only during training, leaving one contextual encoder at retrieval time, according to Perplexity's write-up.
Whole-document chunk embeddings
Independent chunking gives a retriever granular units, but it can remove the heading, definition, alias, or entity that makes a chunk intelligible. Perplexity's technical post describes late chunking as the compromise: encode the document once, then pool representations into separate chunk vectors.
The released model marks chunk boundaries with a learned separator token and mean-pools the token representations inside each chunk. It returns one embedding per chunk, so the index keeps chunk-level granularity while each vector carries information from the rest of its document model card.
The model encodes queries separately from documents. Its card exposes encode_queries for queries and encode for document chunks, reflecting the fixed query and document prefixes used during training model card.
Token-level supervision
Most contextual embedding training marks one chunk as relevant and treats the rest of the document as negative. the gold-chunk explanation identifies the failure mode, while the supporting-context example shows why an answer may need other chunks that define an alias or resolve a reference.
Perplexity's context compression model supplies the richer target:
- It reads the query and positive document jointly and assigns a relevance score to every document token the teacher description.
- The scores are aggregated over arbitrary chunk boundaries, producing a continuous distribution instead of one binary gold label the score-distribution explanation.
- A chunk's target relevance is based on the mean of its top-scoring tokens, then converted into a softmax distribution within the positive document the loss breakdown.
- The student matches that distribution with a forward KL loss, while a document-level InfoNCE loss uses the highest-scoring chunk as the document score the two-objective breakdown.
The context compressor is a training-time teacher. At inference, the released encoder produces the chunk vectors without a second compression or reranking stage, according to Perplexity's method description.
Training recipe
The recipe is broader than the benchmark headline. Perplexity reports roughly 430 public and in-house query-document datasets covering more than 50 languages, with no ConTEB data in training and no context-bench data available during development training details.
The released model uses:
- An in-house 9B-parameter ColBERT retrieval model as its starting point, followed by a projection to 2048 dimensions the model specification.
- Randomly selected chunking strategies during training, with a learned
<|chunk_sep|>token and mean pooling within each chunk the chunking details. - Matryoshka training for 1024 and 2048 dimensions, plus quantization-aware training for native int8 embeddings the embedding formats.
- A model soup that averages several checkpoints from the same training run model construction.
The teacher's token scores can therefore supervise multiple chunking schemes without relabeling the documents. That property is central to the release's training method, not an inference-time feature the flexible-boundary explanation.
Context-bench
ConTEB measures contextual chunk retrieval, but its single-gold-chunk framing does not test whether a system retrieves the evidence that makes an answer verifiable. Perplexity and turbopuffer built context-bench around three tasks: finding the right document among near duplicates, finding an answer chunk whose meaning depends on distant context, and recovering supporting evidence the benchmark scope.
The private benchmark contains 2,099 queries over 38,894 long documents in 21 domains. Perplexity says the model was developed independently of the benchmark and submitted blind, according to the official evaluation description.
Perplexity reports these context-bench results at K=10:
- Answer recall: 45.5% for the preview versus 31.1% for voyage-context-4, a +14.4 point gap the answer-recall result.
- Evidence recall: 40.6% versus 35.6%, a +5.0 point gap the benchmark chart.
- Document recall: 61.6% versus 51.6%, a +10.0 point gap the benchmark chart.
On ConTEB, the preview has the highest average nDCG@10 among the models shown, but the lead is not universal. Perplexity context v1 leads NarrativeQA, while Nemotron-3-8B leads COVID-QA in the comparison chart the ConTEB chart.
Vector formats
The model supports 2048-dimensional output and a 1024-dimensional Matryoshka slice. Its native int8 output changes the storage comparison: Perplexity's chart puts the 1024-dimensional format at 1 KB per vector, versus 8 KB for voyage-context-4 at 2048 dimensions and float32 the vector-size comparison.
At those settings, Perplexity reports that the smaller int8 representation still beats voyage-context-4 on its chunk-retrieval comparison the storage result. The model card says other truncation sizes were not trained model card.
Preview deployment
The weights are publicly available on Hugging Face, but the release remains a preview. The model card says weights, embeddings, and the interface may change without backward compatibility, and it requires transformers>=5.4.0, PyTorch, NumPy, Safetensors, and tqdm; loading uses trust_remote_code=True model card.
Perplexity researcher Antoine Chaffin said the team was still working to make the model available through its API the API status. The card's usage path is therefore local model loading, with normalized embeddings recommended for dot-product scoring or cosine similarity applied to the native unnormalized int8 output model card.
The release also uses two parameter counts across its surfaces: Perplexity's post describes the starting checkpoint as a 9B-parameter ColBERT, while Hugging Face lists the released model size as 8B parameters Perplexity's training post, Hugging Face metadata.