Advertising disclosure: we earn commissions when you shop through the links below.
AMD's Ryzen AI Max PRO 400 Series is a refresh of its large Ryzen AI Max APUs, codenamed Gorgon Halo, built for commercial desktops and workstations that run AI models locally. The chips pair Zen 5 CPU cores with RDNA 3.5 integrated graphics and an XDNA 2 NPU rated at up to 55 TOPS, plus up to 192 GB of unified memory shared by the CPU and GPU. AMD says the platform can run a 300-billion-parameter language model entirely on-device with no cloud connection. Pricing and full retail availability were not announced.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
AMD's Ryzen AI Max PRO 400 Series, codenamed Gorgon Halo, is a mid-generation refresh of the Ryzen AI Max 300 family: the same Strix Halo silicon with a modest clock bump and a much larger memory ceiling. The headline is 192 GB of unified LPDDR5X, up to 160 GB of which can be addressed by the integrated GPU. That single number is what makes this chip interesting for local AI work, because it moves a single x86 client processor into territory that previously required a multi-GPU workstation or a cloud instance.
This is not a data center part. It sits in the prosumer and professional workstation tier, soldered into mini-PCs, mobile workstations, and compact desktops. AMD positions it against Apple Silicon (M4 Max, M3 Ultra) and NVIDIA's GB10-based DGX Spark on the unified-memory side, and against discrete-GPU workstations on the raw compute side. Where it wins is capacity per watt per dollar of desk space: no other x86 client chip lets you hold a 300B-class model resident in memory. Where it loses is bandwidth, and bandwidth is what determines how fast those tokens actually come out.
The Ryzen AI Max PRO 400 lineup announced so far includes three SKUs: the Ryzen AI Max+ Pro 495 (16 cores, 40 RDNA 3.5 CUs), the Max Pro 490, and the Max Pro 485. Non-PRO variants are expected to follow. Pricing and full retail availability were not announced at launch, and systems are expected to ship in the third quarter of 2026.
The specs that matter for inference, in order of how much they constrain you:
For LLM inference, capacity decides whether a model runs and bandwidth decides how fast. The 400 Series attacks the first problem aggressively. 160 GB of GPU-addressable memory comfortably holds:
The 300B claim is real but tight. A 300B dense model at Q4_K_M lands in the 150 GB-170 GB range, leaving very little headroom for KV cache once you push context past a few thousand tokens. Treat 300B as "runs, at low quantization, with modest context" rather than "runs comfortably."
The 256-bit LPDDR5X subsystem is the same class of memory as the 300 Series it replaces. It is roughly an order of magnitude below a discrete card like the RTX 5090, whose GDDR7 delivers well over 1.7 TB/s. That gap shows up directly in single-stream token generation, which is memory-bandwidth-bound for dense models. Compute (the 40 RDNA 3.5 CUs and the XDNA 2 NPU) is not your bottleneck for autoregressive decode; the memory bus is.
Practical implication: dense models feel slow, sparse MoE models feel fast. A 30B MoE activating 3B parameters per token generates tokens several times faster than a 32B dense model on identical hardware, because it reads far less weight per token.
The 55 TOPS XDNA 2 NPU is Copilot+ class silicon. It is excellent for always-on, low-power workloads (small classifiers, speech, vision preprocessing, background agents) and it is what Windows and AMD's AI stack will route small models to. It is not how you run a 70B language model. For anything large, you are on the iGPU via ROCm/HIP or Vulkan through llama.cpp, Ollama, or LM Studio. Plan your software stack around that.
Estimates below assume Q4_K_M quantization unless noted, 160 GB GPU-addressable memory, and current ROCm/Vulkan inference stacks. Treat tokens/sec as ranges, not promises; they move with context length, batch size, and driver maturity.
| Model | Quant | Fits? | Est. single-stream tok/s |
|---|---|---|---|
| Llama 3.1 8B | Q4_K_M | Trivially | 25-40 |
| Mistral 7B / Qwen 2.5 7B | Q4_K_M | Trivially | 25-40 |
| Qwen 2.5 32B | Q4_K_M | Yes | 6-9 |
| Llama 3.3 70B | Q4_K_M | Yes | 4-6 |
| Llama 3.1 70B | Q8_0 | Yes | 3-5 |
| Qwen3 30B-A3B (MoE) | Q4_K_M | Yes | 35-60 |
| gpt-oss-120b (MoE) | Q4_K_M | Yes | 20-40 |
| DeepSeek-R1 671B | Q4 | No (needs ~400 GB) | n/a |
| 300B-class dense | Q4 | Borderline | 1-3 |
Q4_K_M is the quality-to-speed pivot on this hardware, and Q5_K_M is worth the small speed hit if you have memory to spare. Below Q4 you start losing measurable reasoning quality on math and code; above Q6 you are paying bandwidth for gains most users cannot perceive. For MoE models specifically, Q4_K_M is close to free, since only a fraction of weights are read per token.
Vision-language models are a strong fit: Qwen2.5-VL 72B, Llama 3.2 Vision, and Gemma 3 all fit at Q4 with room for image tokens. Image generation (SDXL, Flux) runs comfortably. The real advantage over a 32 GB discrete GPU is context. A 70B model at Q4 with 128K context needs 20-40 GB of KV cache depending on GQA configuration; quantize the KV cache to Q8 or Q4 and you can hold very long contexts without evicting weights. That is the scenario where 192 GB stops being a spec-sheet number and starts being useful.
LoRA and QLoRA fine-tuning of 7B-32B models is feasible. Full fine-tuning of a 70B is not, and pretraining is out of scope entirely. This is an inference and light-adaptation platform.
Local chatbot and assistant users. If you want a private GPT-class assistant with no per-token cost and no data leaving the machine, this is one of the few single-box options that runs 70B models at usable quality and 120B MoE models at genuinely good speed.
Developers building agentic workflows. Multi-agent systems are memory-hungry: several models resident at once (a planner, a coder, a vision model, an embedding model) plus long conversation state. 192 GB lets you keep all of them loaded instead of swapping. For teams building local AI agents in 2026, capacity is often the binding constraint, not FLOPS.
Teams running small inference servers. A mini-PC with a Max+ Pro 495 can serve a handful of concurrent users on a 70B model or a larger MoE. Throughput per box is modest, but so is power draw and noise compared to a GPU server.
Edge and regulated deployment. On-premises inference where data cannot leave the building, or where a cloud dependency is unacceptable, is the clearest fit.
Who should look elsewhere. Anyone whose workload is compute-bound rather than memory-bound: large-batch serving, training, video diffusion at scale, or high-throughput embedding pipelines. A discrete GPU will beat this decisively on raw throughput per dollar.
vs. Apple M4 Max / M3 Ultra. Apple's top silicon offers up to 128 GB (M4 Max) or 512 GB (M3 Ultra) unified memory with substantially higher bandwidth, and a more mature inference stack in MLX. The M3 Ultra configs cost considerably more. Pick Apple if you want maximum bandwidth and are comfortable in macOS. Pick the Ryzen AI Max PRO 400 if you need x86, Windows or Linux, CUDA-adjacent tooling via ROCm, or a specific price point.
vs. NVIDIA DGX Spark (GB10). Similar concept, 128 GB unified, lower capacity than the 400 Series at its 192 GB ceiling. NVIDIA brings CUDA and a far more polished software ecosystem; AMD brings more memory headroom for the largest models. If your stack already assumes CUDA, the switching cost is real.
vs. a discrete GPU workstation (RTX 5090, RTX PRO 6000). A 32 GB RTX 5090 is dramatically faster per token and can run 32B models at high speed, but cannot hold a 70B at Q8 or a 120B MoE at all. The 400 Series trades speed for capacity. If your models fit in 32 GB, buy the GPU. If they do not, this is the cheaper path than a multi-GPU rig.
Bottom line for buyers: choose the Ryzen AI Max PRO 400 Series when model size and context length matter more than tokens per second, and when you want that capability in a single, quiet, low-power box.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.