Nemotron-3-Embed-1B-BF16 is a multilingual text embedding model from NVIDIA for semantic search, retrieval and RAG pipelines. It has about 1.14B parameters, produces 2048-dimension vectors and handles up to 32k tokens of context, with reported support for 34 languages including code retrieval. The model was derived from Ministral-3-3B-Instruct-2512 through two rounds of structured pruning and distillation and is released under the OpenMDW License 1.1 with Apache 2.0 terms. A free NVIDIA NIM endpoint and the BF16 weights on Hugging Face are available, alongside an NVFP4 variant.
A workable 1.14B-parameter dense embedding model from NVIDIA. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 1.5 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 1.5 GB |
| AMD Instinct MI300XAMD | SS | 1.5 GB |
| AMD Instinct MI325XAMD | SS | 1.5 GB |
| AMD Instinct MI355XAMD | SS | 1.5 GB |
Cheapest current cloud rentals with at least 1 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA RTX A4000Vast.ai · On-Demand · 16 GB VRAM | $0.08 |
NVIDIA GeForce RTX 4070Vast.ai · Spot · 12 GB VRAM | $0.09 |
NVIDIA GeForce RTX 5070Vast.ai · Spot · 12 GB VRAM | $0.09 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Nemotron-3-Embed-1B-BF16 is a transformer text embedding encoder built by NVIDIA from the Ministral-3-3B-Instruct-2512 parent through two iterative rounds of structured pruning and distillation. It uses bidirectional attention masking and average pooling over token representations to produce a single 2048-dimension dense vector per input. NVIDIA reports it reaches top retrieval quality among models of comparable size on multilingual benchmarks, and ships a Blackwell-optimized NVFP4 sibling that retains about 99.5% of the BF16 version's reported RTEB score.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every NVIDIA model we track.

Explore the Family
The full Nemotron family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.