Perplexity AI released pplx-embed-v2-late-9B on 2026-10-07 as the larger model in a late-interaction embedding family. It is a ColBERT-style multimodal retriever that embeds text, images and rendered PDF pages into 128-dimensional token vectors compared with MaxSim, so one input cannot mix text and images. The 9B model has 7.4B active parameters, is based on Qwen3.5, and is meant for building document indexes, while a 0.6B sibling shares the same embedding space and can query those indexes. Weights are on Hugging Face under the MIT license and a hosted Perplexity API endpoint is planned but not live.
A situational 9B-parameter dense embedding model from Perplexity AI. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 5.0 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 5.0 GB |
| AMD Instinct MI300XAMD | SS | 5.0 GB |
| AMD Instinct MI325XAMD | SS | 5.0 GB |
| AMD Instinct MI355XAMD | SS | 5.0 GB |
Cheapest current cloud rentals with at least 5 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4060Vast.ai · On-Demand · 8 GB VRAM | $0.08 |
NVIDIA GeForce RTX 3070Vast.ai · Spot · 8 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Perplexity AI released pplx-embed-v2-late-9B as an open-weight, late-interaction embedding model designed for deep document retrieval and visual document indexing. Built on top of a Qwen3.5 backbone with 7.4B active parameters and an overall 9B parameter footprint, the model applies a ColBERT-style multi-vector approach to search. Instead of compressing an entire document or passage into a single dense vector, it maps tokens into individual 128-dimensional vector representations and evaluates relevance using token-level MaxSim operations.
The model marks a shift in visual retrieval architectures by treating rendered PDF pages, images, and raw text as first-class inputs in an end-to-end embedding pipeline. Released under an open-source MIT license, pplx-embed-v2-late-9B is positioned primarily as an indexing workhorse. It shares an identical embedding space with Perplexity's smaller 0.6B companion model, giving engineers the option to build rich offline document indexes using the 9B model while querying them in production with the lightweight 0.6B model.
For local search engines and self-hosted retrieval-augmented generation (RAG) pipelines, this release offers an alternative to multi-stage pipelines that stitch together optical character recognition (OCR), layout parsers, and single-vector embedding models. If your pipeline indexes visually complex documents such as financial reports, charts, or technical manuals, this model handles the structural parsing and semantic indexing directly.
The architecture of pplx-embed-v2-late-9B combines a dense transformer foundation derived from Qwen3.5 with a multi-vector projection layer. While the model contains 9B total parameters, roughly 7.4B parameters are actively engaged during feed-forward inference passes, keeping computation tight for a model in this weight class.
The model processes inputs in discrete modalities. A single input instance can be either text, an image, or a rendered PDF page, but it does not accept mixed or interleaved text-and-image tokens in a single forward pass. Each input is projected into a matrix of token-level embeddings, with each token represented as a 128-dimensional vector.
Relevance scoring relies on the late-interaction MaxSim operator. When scoring a query against a document:
1Score(Q, D) = Sum_{q in Q} Max_{d in D} (q · d)
This late-interaction design preserves granular lexical and positional details that single-vector dense representations discard during mean pooling. Because the 9B model shares its latent geometric space with the 0.6B variant, the vector representations remain fully interoperable. Engineers can index document corpora with pplx-embed-v2-late-9B on dedicated hardware and execute live user search queries using the 0.6B variant on low-power edge machines without retraining or score calibration.
Traditional retrieval setups fail when documents contain complex visual structures like multi-column layouts, figures, tables, and infographic callouts. Running OCR to turn pages into raw text strips spatial hierarchy and visual relationships. pplx-embed-v2-late-9B solves this by embedding visually rendered pages as visual tokens that align directly with natural-language query tokens.
Practical workloads where pplx-embed-v2-late-9B delivers clear advantages include:
Running a multi-vector model locally introduces two resource demands: standard GPU memory for neural network inference, and additional storage and system memory for index storage (since multi-vector indexes scale linearly with the number of tokens in your corpus).
To run pplx-embed-v2-late-9B locally for document embedding and query encoding, target the following GPU VRAM configurations:
For Apple Silicon setups, unified memory simplifies allocation. A Mac Studio or MacBook Pro with an M3 Max or M4 Max and 36 GB or more of unified memory runs full-precision FP16 indexing without paging to disk.
The best quantization for pplx-embed-v2-late-9B in production is Q4_K_M or Q8_0, depending on your throughput needs:
Because pplx-embed-v2-late-9B is an embedding model rather than an autoregressive text generator, performance is measured in indexing throughput (pages or tokens per second) rather than generation tokens per second.
On an RTX 4090:
While standard single-vector embedding models drop directly into standard Ollama workflows, late-interaction models require ColBERT-compatible inference engines or custom PyTorch runtimes to compute MaxSim matrices and handle visual patch inputs.
Engineers can load weights directly via Hugging Face Transformers or integrate with vector stores designed for ColBERT-style tensor storage (such as Vespa, PLAID, or Stanford ColBERT runtime). If you run Ollama as your primary inference runner, look out for community GGUF ports specifically patched for multi-vector late-interaction projection heads.
Evaluating pplx-embed-v2-late-9B requires looking at its performance relative to other vision-capable and multi-vector retrieval engines.
ColPali (built on PaliGemma architectures) popularized visual document retrieval using late interaction.
pplx-embed-v2-late-9B scales up to a 9B base (7.4B active). The larger parameter count gives Perplexity's model deeper semantic understanding over technical text and dense visual layouts.pplx-embed-v2-late-9B stands out because its index is compatible with the 0.6B query encoder, reducing query-side compute costs.BGE-M3 is widely used for enterprise search, supporting dense retrieval, sparse retrieval, and multi-vector ColBERT-style retrieval.
pplx-embed-v2-late-9B eliminates the OCR failure points that hurt BGE-M3 pipelines.For visual document indexing and high-precision RAG deployments running on modern consumer hardware, pplx-embed-v2-late-9B provides an effective balance of precision, architecture scale, and license flexibility.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Perplexity AI model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.