Nvidia released Nemotron-3-Embed-8B-BF16 on July 15, 2026, an open-weight text embedding model for multilingual retrieval, semantic search and RAG. It is an 8B transformer encoder built on Ministral-3-8B-Instruct-2512 with a 32k context window and support for 34 languages plus code retrieval. It ranked first on the multilingual RTEB leaderboard. Licensed under OpenMDW-1.1 and free to download from Hugging Face.
A workable 8B-parameter dense embedding model from Nvidia. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 7.2 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 7.2 GB |
| AMD Instinct MI300XAMD | SS | 7.2 GB |
| AMD Instinct MI325XAMD | SS | 7.2 GB |
| AMD Instinct MI355XAMD | SS | 7.2 GB |
Cheapest current cloud rentals with at least 7 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.04 |
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.06 |
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.07 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Nemotron-3-Embed-8B-BF16 is a text embedding model trained by NVIDIA for retrieval and semantic similarity. It generates dense vector embeddings from multilingual text and is intended as a component of text-based RAG and agentic retrieval systems. The model is a Ministral-3-8B-Instruct-2512 based encoder with bidirectional attention masking, approximately 8B parameters and a hidden size of 4096. It was evaluated across 34 languages and supports code retrieval. It is ready for commercial use and is available on Hugging Face, as an NVIDIA NIM microservice, and via vLLM.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every NVIDIA model we track.

Explore the Family
The full Nemotron family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.