Advertising disclosure: we earn commissions when you shop through the links below.
HP ZGX Fury G1n AI Station is a liquid-cooled 5U-rackable AI workstation from HP built on the NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip. It pairs a 72-core Arm Neoverse V2 Grace CPU with 496 GB of LPDDR5X and a Blackwell Ultra GPU with 252 GB of HBM3e, for 748 GB of coherent memory. HP says it runs models up to 1 trillion parameters locally and it ships with Ubuntu 24.04 LTS plus NVIDIA AI Developer Tools. Pricing and availability were not stated.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
The HP ZGX Fury G1n AI Station is a liquid-cooled, 5U-rackable desktop workstation engineered specifically for massive on-premises generative AI workloads, agentic pipelines, and frontier model deployment. Built around the NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip, this system bridges the gap between traditional enterprise workstations and multi-node data center clusters. By housing 748 GB of unified, coherent memory directly at the desk or in local server racks, it targets organizations requiring absolute data sovereignty, low-latency execution, and deterministic infrastructure costs free from cloud API token fees.
Historically, running frontier-class models locally required multi-GPU server nodes tethered to dedicated server rooms with 3-phase power and loud high-CFM chassis cooling. The ZGX Fury G1n repackages that tier of compute into an acoustically managed, liquid-cooled form factor running Ubuntu 24.04 LTS with preloaded NVIDIA AI Developer Tools. For research labs, enterprise AI teams, and engineers cataloging high-tier systems within an Unknown hardware for AI development registry, the platform stands out as a turnkey execution environment for models that previously could only live in hyperscaler environments.
Deploying the HP ZGX Fury G1n AI Station for AI workloads eliminates the distributed systems overhead typically associated with running multi-node clusters. With a single, unified Grace Blackwell architecture, developers escape complex tensor-parallel networking bottlenecks across PCIe switches, making it a definitive candidate for teams evaluating the best hardware for running AI models locally.
The top models this device can run at 4-bit, ranked by fit and speed.
| Model | Grade | Speed | VRAM |
|---|---|---|---|
| Naive-N0.5-FlashNaiveAI | SS | 42.8 tok/s | 133.5 GB |
| MiMo-V2.6-Flash-RLXiaomi MiMo | SS | 44.2 tok/s | 129.2 GB |
| DeepSeek-V4.1-Flashdeepseek-ai | SS | 41.5 tok/s | 137.8 GB |
| Llama 4 MaverickMeta | SS | 39.1 tok/s | 146.4 GB |
| Llama 3.1 70B InstructMeta |

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.
Generated from this product’s spec sheet. Editor reviews refine it over time.
The compute engine at the core of the ZGX Fury G1n AI Station is the NVIDIA GB300 Grace Blackwell Ultra Desktop Superchip. This SoC integrates a 72-core ARM Neoverse V2 Grace CPU with an NVIDIA Blackwell Ultra GPU across an ultra-high-speed NVLink-C2C (Chip-to-Chip) interconnect delivering 900 GB/s of bidirectional bandwidth. This hardware interconnect allows the CPU and GPU to share a unified physical address space without standard PCIe bus serialization penalties.
1+------------------------------------------------------------------------+2| HP ZGX Fury G1n AI Station Architecture |3| |4| +-----------------------------+ NVLink-C2C +---------------+ |5| | Grace CPU (72-Core) |<===================>| Blackwell GPU | |6| | 496 GB LPDDR5X @ 396 GB/s | (900 GB/s) | 252 GB HBM3e | |7| +-----------------------------+ | @ 7.1 TB/s | |8| +---------------+ |9| Total Coherent Memory: 748 GB |10+------------------------------------------------------------------------+
Memory bandwidth is the primary operational constraint for autoregressive large language model token generation. While conventional multi-GPU workstations relying on DDR5 or GDDR6 top out between 1 TB/s and 2 TB/s aggregate, the 252GB GPU for AI on this system reaches 7.1 TB/s across its HBM3e stack. This level of throughput translates directly into minimal per-token latency during the memory-bandwidth-bound decode phase of LLM inference.
Evaluating HP ZGX Fury G1n AI Station VRAM for large language models highlights its main structural advantage: memory headroom. Because the system provides 252 GB of ultra-fast HBM3e directly on the Blackwell Ultra GPU, coupled with 496 GB of CPU memory via NVLink-C2C, it eliminates standard model-fitting constraints.
1+---------------------------+----------------+--------------------+---------------------+2| Model Family | Precision | Memory Footprint | Execution Tier |3+---------------------------+----------------+--------------------+---------------------+4| Llama 3.1 70B / Qwen 2.5 | FP16 / BF16 | ~140-155 GB | Pure HBM3e (Fast) |5| Llama 3.1 70B | FP8 | ~75-80 GB | Pure HBM3e (Fast) |6| Llama 3.1 405B | FP4 | ~230-240 GB | Pure HBM3e (Fast) |7| Llama 3.1 405B | FP8 | ~430-450 GB | Unified (HBM3e+CPU) |8| DeepSeek-R1 / V3 (671B) | FP4 | ~360-390 GB | Unified (HBM3e+CPU) |9| Trillion-Parameter MoE | FP4 Quantized | ~550-650 GB | Unified (HBM3e+CPU) |10+---------------------------+----------------+--------------------+---------------------+
Models that fit entirely inside the 252 GB HBM3e partition benefit from the uncompromised 7.1 TB/s memory bus:
For architectures exceeding 252 GB, the system leverages its 748 GB coherent memory space. This setup handles hardware for running 1T parameter models using native FP4 quantization pipelines:
The sweet spot for daily engineering tasks is FP8 execution for 70B-class models, granting zero perceptible perplexity loss while leaving massive headroom for multi-turn conversational agents with million-token context buffers.
The ZGX Fury G1n is designed for enterprise engineering environments rather than individual hobbyists:
When deciding on high-end local AI infrastructure, teams typically weigh dedicated Superchips against multi-GPU workstation towers.
A quad-RTX 6000 Ada workstation provides an aggregate 192 GB of VRAM (4x 48 GB GDDR6) over PCIe Gen 4/5 interfaces.
Data center server appliances like the DGX H100 provide raw distributed horsepower (8x SXM GPUs), but require specialized 48-volt, high-density server racks and industrial HVAC installations.
For enterprise teams that require massive parameter scaling without the operational complexity of a data center cluster, the HP ZGX Fury G1n AI Station delivers unmatched local inference performance and architectural flexibility.
| SS |
| 50.7 tok/s |
| 112.8 GB |