Advertising disclosure: we earn commissions when you shop through the links below.
NVIDIA Vera Rubin NVL72 is a rack-scale AI supercomputer from NVIDIA built from 72 Rubin GPUs and 36 Vera CPUs. It targets training and inference for large models, including trillion-parameter models and long-context agentic workloads. NVIDIA lists 20.7 TB of HBM4, 54 TB of LPDDR5X, 75 TB of fast-access memory and a 260 TB/s NVLink domain. NVIDIA claims one-tenth the cost per million tokens and one-fourth the GPUs for MoE training versus GB200 NVL72.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
NVIDIA Vera Rubin NVL72 is a rack-scale AI supercomputer built from 72 Rubin GPUs and 36 Vera CPUs, wired into a single 260 TB/s NVLink domain. It is the successor to the GB200 and GB300 NVL72 systems, and it is purpose-built for the two workloads that broke previous generations' economics: trillion-parameter mixture-of-experts models and long-context agentic inference where a single request can burn tens of thousands of tokens on reasoning and tool calls.
This is not a workstation and not a desktop product. It is a 100% liquid-cooled, third-generation MGX NVL72 rack that needs real facility power, liquid cooling, and rack-scale deployment logistics. The directory entry currently carries no confirmed manufacturer record — if you landed here looking for Unknown hardware for AI development, the platform vendor is NVIDIA. Availability is listed as announced; no MSRP is published. Rack-scale NVL72 systems have historically transacted in the millions of dollars per rack, so plan on rental capacity from a neocloud or hyperscaler as the realistic first touch.
The reason practitioners care about NVIDIA Vera Rubin NVL72 for AI is the efficiency claim. NVIDIA states one-tenth the cost per million tokens and one-fourth the GPUs for MoE training versus GB200 NVL72, plus up to 10x more tokens per megawatt. Those numbers are vendor benchmarks (the cost-per-token figure is measured on Kimi-K2-Thinking at 32K/8K input/output sequence lengths), but they define the tier: this is the top of the market, competing with prior-generation NVL72 racks from NVIDIA itself and non-NVIDIA rack-scale accelerators.
All PFLOPS figures are rack-aggregate. Divide by 72 and you get roughly 50 PFLOPS NVFP4 inference and 35 PFLOPS NVFP4 training per Rubin GPU, with 17.5 PFLOPS per GPU at FP8/FP6 for training. NVFP4 is the headline format here — it's what unlocks both the training and inference numbers, and it's why the efficiency claims land. Video generation, graph neural networks, and physical AI simulation pipelines are compute-bound, not bandwidth-bound, so the raw NVFP4 throughput is what moves those workloads.
The 20.7 TB of HBM4 is not there to hold model weights — it's there to hold KV cache at scale. Weights for even a 1T-parameter model at FP4 land around 500–600 GB, meaning the rack can hold dozens of copies. Run the KV math instead: a 128-layer GQA model with 8 KV heads at 128 head dimension, cached at FP8, costs about 256 KiB per token. A 1M-token context therefore consumes roughly 256 GB. That caps you around 80 concurrent million-token sessions in HBM4 alone, with the 54 TB LPDDR5X tier adding a couple hundred more at a latency penalty when you offload there. Models using MLA-style attention (DeepSeek-R1, DeepSeek-V3) compress the cache far harder and push concurrency much higher.
The 260 TB/s NVLink domain is what allows expert-parallel MoE inference without the all-to-all becoming the bottleneck. Scale-out runs over Quantum-X800 InfiniBand and Spectrum-X Ethernet, with ConnectX-9 SuperNICs and BlueField-4 DPUs handling the data path and offload. For multi-rack DGX SuperPOD deployments, this is the fabric that determines whether a 288-GPU job scales linearly or falls off a cliff.
Model fit is essentially a non-issue. At 20.7 TB of HBM4, the rack holds multiple full copies of any open-weight model in existence. The relevant question is throughput and concurrency, not whether it fits.
Quantization guidance. For serving, NVFP4 is the best quality-to-speed tradeoff on this hardware — the tensor cores are designed around it, and the throughput delta versus FP8 is large. Reserve FP8/FP6 for fine-tuning runs where small quality deltas matter, and BF16 for small-model reference evaluations where you're not throughput-bound.
Tokens per second. Per-stream decode is bandwidth-bound, not compute-bound; a single user on a trillion-parameter MoE will see tens-to-low-hundreds of tokens/sec depending on active parameters. The rack's value is aggregate: thousands of concurrent streams at the same latency budget, which is exactly the shape of agentic traffic. NVIDIA's MLPerf Inference v6.1 debut submission (a preview entry) shows leading system performance at rack scale — check that submission for the workloads closest to yours before sizing.
Multimodal and long-context. Video generation, high-fidelity simulation, and multimodal front-ends all fit; the compute headroom handles the encoder/decoder side without starving the LLM serving path.
vs. NVIDIA GB300 / GB200 NVL72. Same rack footprint, same third-generation MGX design, same NVIDIA software stack — which makes Vera Rubin an in-place upgrade rather than a re-architecture. NVIDIA's own comparison is the strongest case: one-fourth the GPUs for MoE training and one-tenth the cost per million tokens versus GB200 NVL72. If you already operate NVL72 racks, the migration risk is low and the efficiency delta is large. If you're buying your first rack-scale system, there's little reason to buy the outgoing generation.
vs. AMD Helios-class rack systems. The competing rack-scale accelerator approach, typically positioned on power-per-dollar and open rack designs. NVIDIA's durable advantages are the 260 TB/s single NVLink domain and the maturity of the CUDA serving stack — both matter more at 72-GPU scale than on a single card. AMD wins where pricing, supply, or vendor-diversity requirements dominate. Verify current Helios specifications independently; treat vendor-published comparisons from either side as directional.
When to pick the NVL72: you're serving trillion-parameter MoE models at long context, you need expert-parallel all-to-all that doesn't bottleneck, and you can support a liquid-cooled rack. When to look elsewhere: your models fit in a few hundred gigabytes, your traffic is bursty and low-volume, or you'd rather rent capacity than own the cooling loop.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.


Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.

Get a full budget-matched parts list for a local AI workstation.
Mac vs NVIDIA for local inference, if you are still choosing a platform.