On October 7, 2026, Perplexity announced pplx-embed-v2-late, a family of multimodal retrieval models in 0.6B and 9B sizes. Perplexity said both were publicly available on Hugging Face. The models search text and images by comparing individual tokens, rather than compressing each input into a single vector.

How late interaction and MaxSim work

Perplexity’s new embedding models search text and images

An embedding turns an input into numerical vectors that a retrieval system can compare. Many systems represent a whole input with one vector. The late-interaction approach in pplx-embed-v2-late keeps multiple vectors, one for each token, with 128 dimensions per vector.

When a query is matched against a document, each query token finds its strongest match among the document tokens. MaxSim combines those strongest matches into a score for the query-document pair. This preserves token-level matches through scoring, at the cost of storing more vectors per document and doing more work to score candidates than with dense, single-vector retrieval. Perplexity says those costs depend on document length and deployment choices.

A 0.6B model can query an index built with 9B

Perplexity says the two variants share an embedding space. That means a system can use the 9B model to encode documents for an index, then use the 0.6B model to encode search queries against that index. Perplexity says using the smaller model for query encoding can reduce computation at search time.

The scores for this mixed setup are distinct from the symmetric results for either model alone. On ViDoRe v3 image retrieval, Perplexity reports 62.3% nDCG@10 for 0.6B used for both indexing and queries, and 63.5% when 9B indexes documents and 0.6B encodes queries. The model card’s symmetric 9B result is 65.2%.

What the ViDoRe v3 scores measure

Perplexity’s model cards report results for both image retrieval and a separate task using Markdown-converted documents. nDCG@10—normalized discounted cumulative gain at 10—is a ranking-quality measure: it evaluates how relevant results are ordered among the top ten, not the percentage of queries answered correctly.

Model variantViDoRe v3 image nDCG@10ViDoRe v3 Markdown nDCG@10
0.6B62.3%61.2%
9B65.2%64.7%

These are Perplexity-published results for specific benchmark tasks. They give a direct comparison between the two variants on those tasks, not a universal measure of retrieval quality across every dataset or application.

Searching visual documents and handling inputs

Perplexity says the models can retrieve visual-document pages, including rendered PDF pages, without OCR or parsed text. In that workflow, pages are supplied as images. The documented encoding setup uses separate text-only and image-only batches; it does not combine both input types in one batch.

The model cards document native Sentence Transformers integration and list requirements of sentence-transformers >= 6.0.0 and transformers >= 5.4.0. Perplexity also said it would progressively roll out support for late-interaction, dense, and contextual embeddings on its API Platform.