Nvidia released llama-nv-embed-reasoning-3b, a 3.2B parameter text embedding model built on meta-llama/Llama-3.2-3B. It produces dense vectors for retrieval, semantic search and similarity tasks, with a focus on reasoning-heavy content such as multi-step explanations and question-answer pairs. It outputs 3072-dimension embeddings, was trained with a 512 token limit and evaluated at up to 8192 tokens, and supports a 131072 token position range. Weights are published on Hugging Face under a non-commercial Creative Commons license.
A situational 3.2B-parameter dense embedding model from Nvidia. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 5.7 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 5.7 GB |
| AMD Instinct MI300XAMD | SS | 5.7 GB |
| AMD Instinct MI325XAMD | SS | 5.7 GB |
| AMD Instinct MI355XAMD | SS | 5.7 GB |
Cheapest current cloud rentals with at least 6 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.04 |
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.06 |
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.07 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Llama-NV-Embed-Reasoning-3B is Nvidia's 3.2B-parameter dense text embedding model engineered specifically for retrieval tasks involving multi-step logic, technical documentation, and complex question-answer pairs. Built on top of meta-llama/Llama-3.2-3B, the model shifts away from traditional surface-level lexical similarity and shallow bi-encoders. Instead, it leverages a fine-tuned decoder backbone trained with contrastive objectives to map intricate queries and long-form explanatory documents into a shared 3072-dimensional vector space.
Most production embedding models (such as smaller BERT or RoBERTa derivatives) degrade when queries require deducing intermediate premises or synthesizing structured reasoning before identifying a match. By adapting a full 3.2B causal language model into an embedding extractor, Nvidia targets reasoning-aware Retrieval-Augmented Generation (RAG) pipelines where concise, ambiguous queries must align with dense, logically complex target passages. Weights are available on Hugging Face under a non-commercial Creative Commons license, positioning it primarily for research, evaluation pipelines, and internal offline exploration.
For developers aiming to run Llama-NV-Embed-Reasoning-3B locally, its 3.2B dense footprint sits in an operational sweet spot. It provides significantly higher semantic capture than standard 300M-parameter bi-encoders while remaining small enough to run alongside local generator LLMs on single consumer GPUs or unified-memory workstations.
Llama-NV-Embed-Reasoning-3B retains the core dense transformer architecture of Meta's Llama-3.2-3B. Unlike mixture-of-experts (MoE) topologies where active parameters represent a fraction of total capacity, every single forward pass activates all 3.2 billion parameters. This ensures consistent VRAM consumption and uniform computational latency during vector extraction.
meta-llama/Llama-3.2-3BBecause the model is adapted from a decoder-only LLM, it processes text using bidirectional attention masks during the embedding extraction phase, allowing each token representation to attend to both preceding and subsequent tokens. Running the model locally requires transformers (v4.51.0 or newer) along with flash-attn to manage sequence scaling efficiently. While the underlying positional architecture supports up to 131,072 positions, practitioners running local evaluation should keep typical document chunks within 512 to 8,192 tokens to remain within the model's primary training and benchmark distribution.
Llama-NV-Embed-Reasoning-3B is optimized for scenarios where standard cosine similarity fails due to divergent vocabulary between queries and documents. Its primary training focuses on logical synthesis, technical Q&A, and step-by-step rationales.
Note that the weights are governed by a Creative Commons Non-Commercial License (CC-BY-NC-4.0) alongside the Llama 3.2 Community License Agreement. This restricts usage to non-commercial research, local testing, and academic benchmarking.
Embedding generation differs from autoregressive text generation. Instead of generating one token at a time in a memory-bandwidth-bound decode loop, embedding models process the entire sequence in a single forward pass (compute-bound prefill). This yields high compute utilization and makes hardware requirements predictable.
Because the model produces 3072-dimensional vectors across full batches, total VRAM requirements scale with sequence length and batch size:
| Quantization Level | Weight Footprint | Minimum VRAM (Batch 1, 512 tokens) | Recommended VRAM (Batch 8, 4096 tokens) |
|---|---|---|---|
| FP16 / BF16 | ~6.5 GB | 8 GB | 16 GB - 24 GB |
| INT8 / Q8_0 | ~3.5 GB | 6 GB | 12 GB |
| INT4 / Q4_K_M | ~2.1 GB | 4 GB | 8 GB |
Unlike generative models where 4-bit quantization (Q4_K_M) causes minimal degradation in conversational outputs, embedding models are more sensitive to quantization noise in hidden layers.
Embedding performance is measured by batch ingestion speed:
You can run Llama-NV-Embed-Reasoning-3B locally using standard Hugging Face Transformers with FlashAttention:
1pip install transformers==4.51.0 flash-attn==2.6.3 accelerate==0.34.2
For production local microservices, serve the weights through vLLM using its embedding endpoint mode:
1vllm serve nvidia/llama-nv-embed-reasoning-3b --task embed --max-model-len 8192
This launches an OpenAI-compatible /v1/embeddings local server that saturates GPU compute via continuous batching.
Evaluating Llama-NV-Embed-Reasoning-3B requires comparing it against standard small bi-encoders and emerging LLM-based embedders.
Traditional bi-encoders like bge-large-en-v1.5 use roughly 335 million parameters and output 1024-dimensional vectors. BGE-Large is much faster to compute, processes text on basic CPUs, and takes less than 1.5 GB of VRAM. However, BGE-Large struggles with complex, multi-hop reasoning where query and document share little lexical structure. Llama-NV-Embed-Reasoning-3B delivers significantly higher recall on difficult datasets like BRIGHT, but at roughly 10x the parameter overhead and higher compute requirements.
Several practitioners use 7B-class models (like Qwen2-7B or Mistral-7B) fine-tuned for feature extraction. While a 7B embedder offers marginally stronger conceptual alignment, it requires at least 16 GB to 20 GB of VRAM in FP16 and significantly reduces ingestion throughput. Llama-NV-Embed-Reasoning-3B provides a practical middle ground: it captures the reasoning strengths of an LLM backbone while remaining compact enough (3.2B parameters) to fit onto affordable hardware alongside a primary generation model.
For local AI systems in 2026 handling complex, reasoning-intensive knowledge bases, Llama-NV-Embed-Reasoning-3B stands out as a specialized, highly capable offline embedding model, provided your application fits within its non-commercial licensing terms.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every NVIDIA model we track.

Explore the Family
The full Llama family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.