Advertising disclosure: we earn commissions when you shop through the links below.
AMD's Ryzen AI Max+ PRO 495 is a Gorgon Halo client processor built for local AI workstations and business PCs. It has 16 Zen 5 CPU cores clocked up to 5.2 GHz, Radeon 8065S graphics with 40 RDNA 3.5 compute units, and a 55 TOPS XDNA 2 NPU. Partner systems pair it with up to 192GB of unified LPDDR5X memory, with up to 160GB assignable to the GPU, which AMD says allows running models above 300B parameters at 4-bit quantization without cloud offload. One of the first systems, Minisforum's MS-S1 MAX-P495, is priced at $7,399 in the US.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
The AMD Ryzen AI Max+ PRO 495 is a high-end mobile and compact workstation APU from the Gorgon Halo architecture family, built on TSMC's 4 nm process. Designed to deliver workstation-class local inference in power-constrained desktop and mobile form factors, the processor integrates 16 Zen 5 CPU cores (32 threads) clocked up to 5.2 GHz, 64 MB of L3 cache, a 40 compute unit Radeon 8065S GPU on the RDNA 3.5 architecture, and a dedicated 55 INT8 TOPS XDNA 2 NPU.
What sets the AMD Ryzen AI Max+ PRO 495 for AI apart from conventional workstation processors is its memory architecture. Partner systems like the Minisforum MS-S1 MAX-P495 route up to 192GB of unified LPDDR5X memory directly to the APU, allowing up to 160GB to be allocated dynamically to the GPU. This eliminates the standard PCIe transfer bottleneck and provides a single, contiguous pool of VRAM capable of fitting frontier-class open models entirely on-chip.
For engineers evaluating the best hardware for running AI models locally, this processor targets the gap between standard consumer desktop GPUs (which max out at 24GB VRAM) and expensive enterprise data center accelerators. By delivering a workable 160GB GPU for AI in a thermal envelope ranging from 45 W to 120 W (55 W default TDP), the APU provides local execution for multi-agent systems and massive context windows without requiring high-voltage multi-GPU server infrastructure.
The viability of the AMD Ryzen AI Max+ PRO 495 AI inference performance comes down to three factors: usable memory footprint, memory throughput, and compute efficiency across its heterogeneous execution units.
1+------------------------------------------------------------------------+2| AMD Ryzen AI Max+ PRO 495 Architectural Overview |3+------------------------------------------------------------------------+4| CPU: 16 Zen 5 Cores / 32 Threads @ up to 5.2 GHz (64 MB L3 Cache) |5| GPU: Radeon 8065S (40 RDNA 3.5 Compute Units) |6| NPU: XDNA 2 Neural Processing Unit (55 INT8 TOPS) |7| System Memory: Up to 192GB LPDDR5X-8000 / LPDDR5X-8533 |8| Max VRAM Allocation: Up to 160GB Unified VRAM |9| Configurable TDP: 45 W to 120 W (55 W Base) |10+------------------------------------------------------------------------+
Because autoregressive token generation is memory-bandwidth bound, the APU's LPDDR5X bus dictates prompt-to-token velocity. While a discrete desktop card like the RTX 4090 offers higher raw memory bandwidth (1,008 GB/s), it is hard-capped at 24GB of VRAM. The Ryzen AI Max+ PRO 495 prioritizes total parameter footprint over raw bandwidth, giving engineers the ability to load model sizes that would otherwise require multiple discrete graphics cards or data center rentals.
The AMD Ryzen AI Max+ PRO 495 VRAM for large language models unlocks weights that have historically been out of reach on client hardware. It stands as dedicated hardware for running 300B+ parameter models on a single compact system.
The sweet spot on the Ryzen AI Max+ PRO 495 local LLM runtime is Q4_K_M to Q8_0 for 70B parameter models. At Q4_K_M, a 70B model maintains near-lossless perplexity relative to FP16 while drastically reducing the memory read payload per token.
In terms of AMD Ryzen AI Max+ PRO 495 tokens per second:
The APU is built directly for long-running, autonomous agents. As teams adopt multi-agent frameworks like LangGraph, AutoGen, and CrewAI, local systems need to hold both the agent logic and active working memory. The 160GB allocatable memory pool allows developers to deploy an agent team locally, with a 70B model serving as the orchestrator and multiple 8B specialist models running in parallel. This makes it a strong contender for the best hardware for local AI agents 2026.
Organizations handling proprietary codebases, patient records, or financial disclosures cannot route prompts to public cloud APIs. Pre-built systems from hardware integrators provide turnkey, unknown hardware for AI development setups that fit under a desk or in a server rack, operating fully air-gapped without recurring cloud inference subscriptions.
While primary pre-training requires distributed clusters, the Ryzen AI Max+ PRO 495 can execute parameter-efficient fine-tuning (PEFT, QLoRA) on 70B models locally. With 160GB of addressable space, optimizer states and gradient checkpoints can be allocated directly in memory alongside the base model weights.
1+---------------------------------------------------------------------------------------+2| Comparison: High-Capacity Local Inference Platforms |3+-----------------------------------+--------------------+--------------+---------------+4| Platform | Usable VRAM | Memory Band. | Typical TDP |5+-----------------------------------+--------------------+--------------+---------------+6| AMD Ryzen AI Max+ PRO 495 | Up to 160GB | ~260 GB/s | 45W - 120W |7| Apple M4 Max (128GB Unified) | ~96GB - 100GB | ~546 GB/s | 45W - 100W |8| 2x Nvidia GeForce RTX 4090 | 48GB (24GB x 2) | 2,016 GB/s | 600W - 900W |9| 1x Nvidia RTX 6000 Ada Generation | 48GB | 960 GB/s | 300W |10+-----------------------------------+--------------------+--------------+---------------+
Apple's unified memory architecture has led consumer-accessible local LLM inference for several years. High-tier M-series chips offer higher memory bandwidth (up to 800+ GB/s on Ultra models), resulting in faster raw tokens per second on models up to 70B.
However, the Ryzen AI Max+ PRO 495 provides a higher addressable ceiling than standard 128GB configurations, allocating up to 160GB directly to compute on an open x86-64 platform. For developers running containerized Linux stacks, native ROCm workflows, and standard x86 toolchains, the Ryzen AI platform avoids macOS virtualization overhead and Metal framework compatibility limits.
A workstation equipped with two Nvidia RTX 4090 cards offers 48GB of total VRAM and superior compute performance, but it cannot load a 70B model at FP16 or a 300B+ model at any practical quantization without heavy offloading. Furthermore, dual 4090 configurations draw between 600 W and 900 W under load, requiring dedicated 1200 W power supplies and specialized cooling.
The Ryzen AI Max+ PRO 495 sacrifices the raw token speed of discrete GDDR6X/HBM memory, but it consumes a fraction of the power and provides more than triple the VRAM footprint of a dual-4090 system. For teams whose bottleneck is model capacity rather than lightning-fast token output, the 495 serves as an efficient, highly capable best AI chip for local deployment.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.