Cohere launched Embed 5 Pro on September 30, 2026. It is an enterprise embedding model with a 128K-token context window, multimodal inputs, and support for over 100 languages. It uses Matryoshka embeddings and supports float, int8, and binary formats. Pricing is $0.12 per million text tokens and $0.40 per million image tokens. It shares an embedding space with Embed 5 Fast, so indexes built with Pro can be queried with Fast.
A workable dense embedding model from Cohere. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 0.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 0.5 GB |
| AMD Instinct MI300XAMD | SS | 0.5 GB |
| AMD Instinct MI325XAMD | SS | 0.5 GB |
| AMD Instinct MI355XAMD | SS | 0.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4070Vast.ai · Spot · 12 GB VRAM | $0.08 |
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.09 |
NVIDIA RTX A4000Vast.ai · On-Demand · 16 GB VRAM | $0.09 |
NVIDIA GeForce RTX 5060Vast.ai · Spot · 8 GB VRAM | $0.09 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cohere launched Embed 5 Pro on September 30, 2026 as the high-quality tier of its Embed 5 family. It is an enterprise embedding model: text and images go in, vectors come out. There is no generative output, no chat template, and no reasoning trace. It exists to make retrieval work, which means RAG pipelines, semantic search, agentic retrieval loops, and the kind of document-heavy enterprise workloads where the quality of the vector determines whether the right passage reaches the language model at all.
The specs that matter: a 128,000-token context window, a dense architecture with an undisclosed parameter count, Matryoshka embeddings with selectable output dimensions of 256, 512, 768, 1024, 1536, or 2048, and native support for float, int8, and binary embedding formats. It handles more than 100 languages and accepts image inputs alongside text. Pricing through the Cohere API is $0.12 per million text tokens and $0.40 per million image tokens.
The most consequential design decision is not the context window. It is that Embed 5 Pro shares an embedding space with Embed 5 Fast, the cheaper tier at $0.08 per million text tokens. You can index a corpus with Pro for maximum recall, then serve live queries with Fast against the same vector database, no reindexing required. For teams running large retrieval systems, that decouples indexing quality from query latency and cost, which is a tradeoff most embedding stacks force you to make once and live with.
Embed 5 Pro is dense, meaning every parameter is active on every forward pass. There is no mixture-of-experts routing, so latency is predictable and there is no expert-load imbalance to tune. Cohere has not published the parameter count, which is normal for a hosted embedding model but does mean you cannot size hardware from a spec sheet. You size it from measurements.
The 128K-token context window is the headline capability. Most embedding models cap out between 512 and 8K tokens, which forces chunking, and chunking is where retrieval quality usually dies: tables get split from their headers, a clause gets separated from the sentence that qualifies it, and a chart loses its caption. At 128K you can embed a full contract, a complete technical manual section, or a multi-page financial filing as a single vector. That does not remove the need for chunking in large corpora, but it makes chunk boundaries a tuning choice rather than a hard constraint.
Matryoshka embeddings let you truncate the output vector and keep most of its retrieval quality. A 2048-dimension embedding truncated to 256 dimensions is still usable, which cuts index storage by 8x and speeds up similarity search proportionally. The int8 and binary output formats extend that further: binary embeddings are typically used as a first-pass filter, with full-precision vectors used to rerank the shortlist. In practice, most production systems will store binary or int8 vectors for the candidate search and keep float vectors for rescoring.
The undisclosed parameter count matters less than it sounds. Embedding models are far smaller than generative models of comparable context length, and the memory profile is dominated by activations over long sequences, not by weights.
It is not a reranker, not a classifier, and not a generative model. If you need cross-encoder reranking, that is a separate component in the stack.
Be clear about what "locally" means here. Cohere has not published open weights for Embed 5 Pro, and the license is unspecified. There is no GGUF, no Ollama tag, and no Hugging Face checkpoint. The only local path is self-hosting the model on your own infrastructure. Cohere states that both Embed 5 Pro and Embed 5 Fast can be self-hosted with vLLM, and the models are also available through Cohere Model Vault, Microsoft Foundry, and Amazon SageMaker for private deployment.
That said, self-hosted embedding inference is one of the more forgiving local workloads, and the hardware picture is manageable:
On quantization: int8 and binary are native output formats, not just storage tricks, so you get most of the compression benefit without touching the weights. If you need to shrink the weights themselves, int8 is the safe floor. Aggressive 4-bit weight quantization is a bad trade for an embedding model: retrieval recall degrades measurably on MTEB-style benchmarks, and unlike a chat model, you cannot hear the degradation in the output. You just silently retrieve worse documents. There is no Q4_K_M path here anyway, since no GGUF weights exist.
On throughput: embedding models are measured in tokens ingested per second, not generated tokens per second. On a 24GB card with 512-token chunks, expect low thousands of tokens per second per stream, scaling with batch size. Long 128K sequences drop sharply because attention cost is quadratic in sequence length. Budget your index build accordingly.
Ollama is the fastest way to get local embeddings running in general, and nomic-embed-text or bge-m3 will have you producing vectors in under five minutes. It will not run Embed 5 Pro. If your requirement is genuinely "runs in Ollama," this model is not the answer.
Embed 5 Pro vs. Cohere Embed 4. Embed 4 is the previous generation, also 128K context, priced at $0.12 per million text tokens and $0.47 per million image tokens. Embed 5 Pro is cheaper on image tokens and adds the shared embedding space with Embed 5 Fast, which Embed 4 does not offer. If you are already on Embed 4, the migration case is the Pro/Fast index-sharing model plus the image token price cut.
Embed 5 Pro vs. Voyage 4 Large. Voyage 4 Large is priced at $0.12 per million tokens, the same as Embed 5 Pro on text, but its context window is 32K. If your documents fit in 32K and you are not indexing images, Voyage 4 Large is a legitimate alternative and worth benchmarking on your own data. If you need 128K documents, multilingual coverage, or image retrieval, Embed 5 Pro is the stronger fit.
Embed 5 Pro vs. open-weight embedding models. BGE-M3 and similar open-weight models run entirely on your hardware with no per-token cost and no vendor dependency. For high-volume, English-dominant, short-chunk retrieval, they are often good enough, and the cost difference at scale is substantial. Choose Embed 5 Pro when multilingual quality, 128K context, image inputs, or the Pro/Fast index-sharing architecture actually move your retrieval metrics. Benchmark both on your corpus before committing; embedding quality differences that look small on a leaderboard frequently show up as large differences in end-to-end RAG answer accuracy.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Cohere model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.