Perplexity Research and turbopuffer unveiled pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG. Each chunk is encoded with the full document in view. The key change is the training signal: the model learns to retrieve an answer plus the context needed to verify it. Not a single “gold” passage, but supporting fragments together. Want to know if you can run it now, how training works, and what benchmarks show? Let’s walk through it.
What is pplx-embed-v2-context-9b-preview and how is it different?
It is a contextual embedding model for RAG where each chunk sees the entire document. The model learns to retrieve the answer together with evidence, rather than chase a single “gold” snippet. This aims to reduce fact-checking errors and better capture cross-sentence dependencies. It shifts the training signal and lets each chunk vector carry document-level context.
Standard RAG splits long texts into chunks. Definitions, headings, or entities often sit outside an isolated chunk. Contextual embeddings address this with late chunking: the document is encoded in one pass and then pooled per chunk. This preserves relationships that would otherwise be lost during hard splitting. The model encodes the whole and only then “cuts,” keeping dependencies inside vectors.
The main shift lies in training. Instead of labeling one chunk as “gold,” the model receives a signal to retrieve the answer plus context sufficient to verify it. That avoids punishing helpful sentences that fall outside the single labeled window. The system becomes less brittle to labeling choices and chunk boundaries. Retrieval reflects how evidence really appears across documents.
In practice, you build the usual chunk index, but vectors already encode relevant interdependencies. Index cost still depends on vector size, like in classic setups. The idea is straightforward: if a chunk carries document context, the odds of surfacing the right clue rise, and answer verification gets more reliable. You gain robustness without adding inference overhead.
Can you deploy it today and what does setup require?
Yes. The model ships as a self-hosted preview with weights on Hugging Face under the MIT license. Loading requires transformers version 5.4.0 or newer, with trust_remote_code=True. It is not yet available via the Perplexity API. The model card notes that weights and interface may change without backward compatibility. Ready to test it locally and control your environment?
This mode gives you full autonomy. You can integrate it into your RAG pipeline and experiment with indexing strategies. Upgrading transformers is a required step. The trust_remote_code flag allows the repository’s custom logic to load correctly at runtime.
Keep the preview status in mind. Since this is a preview, the authors explicitly warn that weights and interfaces may change without backward compatibility. Plan index rebuilds and dependency updates so you can pivot production quickly if needed. That cadence is common during fast-moving research iterations.
There is no Perplexity API access for now. That means cloud integrations should wait or be implemented on your own. If you have strict latency and data control needs, self-hosting is actually an advantage. You avoid network dependency and keep full control over resources and privacy.
Why does the single gold passage fall short in RAG?
Because real answers live in context, not in isolated fragments. When training labels one “gold” chunk, all others become “negatives.” Those include sentences that make the answer checkable. The scheme simplifies reality too much. It ignores that a chunk depends on entities, headings, and definitions found elsewhere in the document.
Contextual models fix this with late chunking. The document is encoded once, then chunk representations are formed by pooling. Embeddings keep dependencies that would be lost under rigid boundaries. The model sees the whole first and only then “slices,” retaining relationships inside vectors. That preserves evidence ties during retrieval.
Perplexity names three more issues with classic labeling. First, binary labels are coarse and miss degrees of usefulness. Second, LLM annotation cost grows linearly with dataset size. Third, labels are bound to a single chunking strategy and transfer poorly to others. Want flexibility across boundaries? You need a different signal.
So the new model learns to pull the answer along with context sufficient for verification. That naturally reduces penalties for helpful fragments that missed the lone “gold” chunk. In real documents, this addresses critical cases: table references, term definitions, or linking names and pronouns across paragraphs. The result is fewer failures at the verification step.
How does training work and what is the model architecture?
A teacher leads training: Perplexity’s query-aware context compression model. It reads the query and document together and scores every token. Chunk relevance is the mean of the top-n token scores inside each chunk. A “soft” target is then formed: a temperature-scaled softmax over chunks in the positive document, with zero for chunks in other documents. The student matches the teacher via forward KL divergence.
A document-level InfoNCE loss runs in parallel. A document scores as its best chunk, inspired by MaxSim in ColBERT. Each batch samples a random chunking strategy. Chunks are separated by a learned token <|chunk_sep|> and mean-pooled. The teacher runs only during training, so inference adds no latency or storage. Flexible chunk boundaries are preserved without re-annotation.
The model starts from an in-house 9B ColBERT retrieval model. A linear projection outputs 2048-dimensional vectors. Matryoshka training additionally supports 1024 dimensions. Quantization-aware training enables native int8 embeddings. The release is a “model soup” of several checkpoints. Training used roughly 430 datasets across over 50 languages, with no ConTEB data. The design blends efficiency with multilingual coverage.
An interactive explainer illustrates why isolated chunks often fail. The teacher scores tokens, those scores are aggregated per chunk, and a soft target emerges. Changing chunk boundaries needs no fresh labels, since token scores simply re-aggregate. The example is modeled on a lease scenario, with illustrative scores showing the path from token scoring to the student’s KL alignment. It makes the distillation steps tangible.
“A token-level teacher replaces the single gold-chunk label.”
What do the benchmarks show and how does it compare?
On context-bench at K = 10, the model reports 45.5% Answer@10, 40.6% Evidence@10, and 31.1% All-Evidence@10. Document recall reaches 15.2% at K = 1 and 61.6% at K = 10. Versus voyage-context-4, it leads by 14.4 points on Answer@10 and by 5.0 points on Evidence@10. Each model is ranked exhaustively against all chunks, so index settings play no role. This isolates retrieval quality cleanly.
On ConTEB, it shows the highest average nDCG@10 among the models displayed. pplx-embed-context-v1-4B wins NarrativeQA, and Nemotron-3-Embed-8B wins COVID-QA. In general retrieval, the model has the best average on query-to-chunk tasks and trails voyage-context-4 slightly on query-to-document. That underscores its focus on contextual chunk retrieval.
For storage, contextual embeddings use one vector per chunk, same as a normal chunk index. Cost depends on vector size. Perplexity reports that 1024-dim int8 (1 KB per vector) slightly exceeds voyage-context-4 at 2048-dim float32 (8 KB) on its chunk-retrieval suite. Chunk-size sensitivity shifts the average from 81.0% to 79.9% between 64 and 512 tokens. The formula is simple: bytes = dimensions × bytes per value.
Feature-wise, pplx-embed-v2-context-9b-preview offers contextual chunk embeddings with open MIT weights and self-hosting; voyage-context-4 provides contextual embeddings as a hosted API (Voyage, MongoDB Atlas) at $0.12 per 1M tokens with the first 200M free; pplx-embed-context-v1-4B is open-weight MIT with Perplexity API access; Nemotron-3-Embed-8B ships open weights under OpenMDW-1.1 and uses independent per-chunk embeddings. Dimensions, quantization, and context windows differ across them.
Parameters: 9B per the blog (Hugging Face lists 8B), 4B for pplx-embed-context-v1, and about 8B for Nemotron-3-Embed-8B. Supported dimensions: 2048 and 1024 for pplx-embed-v2-context-9b-preview; 2048, 1024, 512, 256 for voyage-context-4; 2560 Matryoshka for pplx-embed-context-v1-4B; 4096 sliceable for Nemotron. Quantization: native int8 for pplx-embed-v2-context-9b-preview, multiple int8/uint8/binary variants for Voyage, int8/binary for pplx-embed-context-v1-4B, and float for Nemotron.
Context windows: pplx-embed-v2-context-9b-preview is evaluated up to 32,768 tokens; voyage-context-4 offers 32K per request and 120K with auto-chunking; pplx-embed-context-v1-4B supports 32K; Nemotron-3-Embed-8B supports 32,768. Auto-chunking exists for voyage-context-4; the others do not include it. In terms of access: either self-host or API, depending on the model. If you need full autonomy and MIT terms, the new preview slots neatly into your stack.
Based on Perplexity Research and turbopuffer.