Advertising disclosure: we earn commissions when you shop through the links below.
INNO3D's AWS-511Z-R1 is a full tower workstation aimed at AI inference, engineering simulation, 3D rendering and digital twin work. It takes one Intel Xeon W-3400 or W-3500 processor, up to 2 TB of DDR5 memory, and either four NVIDIA RTX PRO 4500 Blackwell Workstation Edition GPUs or two RTX PRO 6000 Blackwell Workstation Edition GPUs. Storage covers three M.2 NVMe slots plus up to 12 front drive bays, with 10GbE and 1GbE networking and dual redundant 3000 W 80 Plus Titanium power supplies. GPUs are sold separately and INNO3D has not published pricing.
The INNO3D Tower Workstation AWS-511Z-R1 is a full tower, single-socket Xeon W platform designed to hold up to four workstation-class Blackwell GPUs. It sits in the professional workstation tier, above consumer desktops and below rack-mounted data center servers, and it exists for one reason: putting 96 GB to 192 GB of GPU memory on or under a desk without renting cloud capacity.
INNO3D is best known for graphics cards, and this machine reads like a chassis built around that expertise. The W790 board takes one Intel Xeon W-2400, W-2500, W-3400, or W-3500 processor on LGA4677, with 16 DIMM slots and a 2 TB DDR5 3DS RDIMM ceiling on the W-3400/W-3500 parts (1 TB on W-2400/W-2500). Six PCIe Gen 5 x16 slots (one wired at x8) mean the GPU layout is a real choice rather than a fixed configuration: four RTX PRO 4500 Blackwell Workstation Edition cards at 32 GB each, or two RTX PRO 6000 Blackwell Workstation Edition cards at 96 GB each.
For anyone hunting for the best hardware for running AI models locally, the interesting number here is not core count. It is VRAM. A 70B parameter model at Q4 needs roughly 40 GB. At FP8 it needs about 70 GB. At FP16 it needs 140 GB. The AWS-511Z-R1 is one of the few tower platforms that clears all three of those bars without offloading weights to system RAM, where token generation collapses to single-digit throughput.
Inference speed on a dense transformer is memory-bandwidth bound, not compute bound. Every token generated requires reading the full weight set once, so tokens per second scales roughly with bandwidth divided by model size.
GPUs are sold separately and INNO3D has not published pricing or a system-level MSRP. Note also that INNO3D's own product literature states no GPU or PSU is included with the base product, so confirm what the reseller actually ships before ordering.
Estimates below assume batch size 1, a modern runtime (vLLM, TensorRT-LLM, or llama.cpp), and reasonable context lengths. Real numbers vary with quantization format, KV cache pressure, and concurrency.
Dual RTX PRO 6000 (192 GB)
The 192 GB configuration is the point where a local LLM stops feeling like a demo. With 70B at FP8 you retain well over 100 GB for KV cache, which is what enables long-context work and multiple concurrent agent sessions without eviction thrash.
4 x RTX PRO 4500 (128 GB)
Quantization guidance: on Blackwell, FP8 is the quality-to-speed sweet spot for models that fit. It is natively supported, halves memory versus FP16, and costs far less accuracy than 4-bit. Drop to Q4_K_M or AWQ only when you need the headroom for context or concurrency.
Dell Precision 7960 Tower: dual-socket Xeon Scalable, up to four GPUs, similar memory ceiling. Dell offers better enterprise support, validated driver stacks, and global service contracts. INNO3D counters with a single-socket design that avoids NUMA complexity and a 3000 W redundant power configuration that is unusual in a tower. Pick Dell if procurement and support matter more than raw configuration flexibility.
A self-built 4 x RTX 5090 rig: 32 GB per card gives 128 GB total for substantially less money, and the consumer cards have higher per-GPU bandwidth than the RTX PRO 4500. What you give up is ECC memory, blower-style cooling designed for multi-GPU density, ISV certification, a BMC, and any vendor warranty covering the system as a whole. For a lab or a startup that can tolerate that, it is the cheaper path to 128 GB.
Choose the AWS-511Z-R1 when you need 192 GB of VRAM in a single tower with redundant power, remote management, and a validated chassis that will not thermal-throttle four cards under sustained load. If your models fit in 48 GB, this is more machine than you need.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.