Advertising disclosure: we earn commissions when you shop through the links below.
AMD's Ryzen AI Halo is a compact developer platform for running large language models and agentic AI locally. It pairs a 16-core Ryzen AI Max+ 395 (Strix Halo, Zen 5) with Radeon 8060S graphics, a 50 TOPS XDNA 2 NPU and 128GB of unified LPDDR5X-8000 memory. AMD says it runs models of up to 200B parameters and ships with a 2TB NVMe SSD, Wi-Fi 7, 10GbE and a choice of Windows 11 Pro or Linux. It is priced at $3,999 and pre-orders opened in June 2026.
Manufacturer's suggested retail price. Current prices can be higher or lower. This is not a live price.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
The Ryzen AI Halo is a compact, high-capacity workstation designed specifically for running massive local language models and complex agentic pipelines without relying on cloud infrastructure. Built around AMD's Strix Halo architecture and cataloged under Unknown hardware for AI development, the system packages a 16-core Ryzen AI Max+ 395 processor, integrated Radeon 8060S graphics, and an expansive 128GB of unified LPDDR5X memory into a small-form-factor chassis. At an MSRP of $3,999, it directly targets AI researchers, software engineers, and privacy-conscious teams that require local execution for architectures spanning up to 200 billion parameters.
In the current desktop hierarchy, the system sits in the prosumer and professional edge tier. It bridges the gap between consumer GPUs constrained by 16GB to 24GB of VRAM and expensive multi-GPU enterprise racks. While dual-GPU setups often require high power draws, complex PCIe lane bifurcation, and large cases, the Ryzen AI Halo brings unified memory pooling into a one-box desktop footprint. For developers building multi-agent workflows, evaluating local code assistants, or hosting self-contained inference APIs, it serves as a dedicated local server with out-of-the-box support for both Linux and Windows 11 Pro.
Selecting the best hardware for running AI models locally usually means balancing raw compute throughput against total addressable memory. The Ryzen AI Halo prioritizes capacity above all else. By granting the graphics compute units direct, low-latency access to the entire 128GB system memory pool, it fundamentally eliminates the VRAM ceiling that prevents standard desktop machines from loading modern high-parameter models.
The core of the Ryzen AI Halo is the Ryzen AI Max+ 395 APU. It combines 16 Zen 5 CPU cores (32 threads running at a 3.0GHz base and up to 5.1GHz boost) with a Radeon 8060S graphics engine sporting 40 compute cores clocked up to 2.9GHz. Supplementing the CPU and GPU is AMD's XDNA 2 neural processing unit, delivering 50 TOPS of INT8 compute for low-power background inference, voice pipelines, and continuous sensor processing.
1Hardware Architecture Overview:2+-------------------------------------------------------------+3| Ryzen AI Max+ 395 |4| +--------------------+ +---------------+ +------------+ |5| | Zen 5 CPU (16c/32t)| | Radeon 8060S | | XDNA 2 | |6| | 3.0 GHz / 5.1 GHz | | 40 CUs @2.9GHz| | (50 TOPS) | |7| +--------------------+ +---------------+ +------------+ |8| | | |9| +-------------------------------------------------------+ |10| | 128GB Unified LPDDR5X-8000 Memory (256-bit Bus) | |11| | Bandwidth: 256 GB/s | |12+--+-------------------------------------------------------+--+
For practitioners evaluating Ryzen AI Halo AI inference performance, memory bandwidth is the primary operational metric. Autoregressive token generation is memory-bandwidth bound: each token generated requires every active weight in the model to be read from memory into compute cores. The system uses a 256-bit quad-channel memory bus configured with LPDDR5X-8000, yielding a theoretical peak memory bandwidth of 256 GB/s.
While 256 GB/s is lower than the multi-terabyte bandwidth found on high-end enterprise accelerators like the NVIDIA H100 or Apple M-series Ultra chips, it is paired with a massive 128GB capacity. This allows the system to act as a 128GB GPU for AI without requiring model splitting across PCIe slots. Prompt evaluation (prefill) scales across the 40 RDNA 3.5 compute units, while the Zen 5 CPU cores handle orchestration, vector database querying, and tool-calling code execution natively.
The platform includes a 2TB NVMe SSD for local model weight storage, 10GbE onboard networking for high-speed local inference serving, Wi-Fi 7, and four USB-C ports. The integrated 10GbE interface ensures that when deployed as a headless inference node, network latency does not bottleneck client streaming.
The main justification for deploying a Ryzen AI Halo local LLM workstation is its ability to fit architectures that outright crash standard consumer cards. With 128GB of unified memory, users can run models ranging from lightweight 8B configurations up to heavily quantized 200B parameter models.
1Model Size & Quantization Fit (128GB Memory Pool):2+----------------------+--------------+-------------+----------------------+3| Model Architecture | Quantization | Memory Size | Context Window Space |4+----------------------+--------------+-------------+----------------------+5| Llama 3.1 8B | FP16 / Q8_0 | 8GB - 16GB | Full 128k Context |6| Qwen 2.5 32B | Q8_0 / Q6_K | 26GB - 34GB | Full 64k Context |7| Llama 3.1 70B | Q4_K_M | ~42GB | Up to 64k Context |8| Llama 3.1 70B | Q8_0 | ~75GB | Up to 32k Context |9| Mixtral 8x22B (MoE) | Q4_K_M | ~80GB | 32k Context |10| Command R+ 104B | Q4_K_M | ~65GB | 32k Context |11| Dense 200B Class | Q3_K_M / Q2 | 90GB - 110GB| Limited (8k-16k) |12+----------------------+--------------+-------------+----------------------+
Dense models such as Llama 3.1 8B, Gemma 2 9B, and Qwen 2.5 14B or 32B run with exceptional comfort. At 8-bit or unquantized 16-bit precisions, these models consume only a fraction of the 128GB memory pool.
Because the weights occupy less than 35GB of memory, the remaining 90GB+ can be allocated entirely to the KV cache, enabling expansive context windows up to the architectural maximum of 128k tokens without running out of memory.
For complex coding and reasoning, 70B models represent the primary operational tier:
While 5 tokens per second is not conversational speech speed, it is completely practical for background agentic operations, synthetic data generation, and complex code reviews. Quantizing down to Q3_K_M increases throughput slightly, but Q4_K_M remains the sweet spot for preserving benchmark performance and reasoning stability.
The Ryzen AI Halo is explicitly qualified as hardware for running 200B parameter models. Dense models in the 120B to 200B range fit when using modern 3-bit or 4-bit quantizations (such as Q3_K_M or IQ3_M). Furthermore, Mixture-of-Experts (MoE) architectures excel here:
The machine is engineered for professionals whose workflows are constrained by memory capacity rather than absolute generation latency.
llama.cpp or vLLM.When assessing the Ryzen AI Halo for AI deployments, it is best compared against the NVIDIA DGX Spark (GB10) and high-tier Apple Mac Studio configurations.
1Competitive Landscape:2+------------------------+-------------------+------------------+-------------------+3| Feature | Ryzen AI Halo | NVIDIA DGX Spark | Apple Mac Studio |4+------------------------+-------------------+------------------+-------------------+5| Memory Capacity | 128GB Unified | 128GB Unified | 128GB - 192GB |6| Memory Bandwidth | 256 GB/s | ~500+ GB/s | 800 GB/s |7| Architecture | x86_64 (Zen 5) | ARM (Grace) | ARM (Apple M) |8| Operating System | Linux / Win 11 | Linux (Ubuntu) | macOS |9| MSRP | $3,999 | ~$4,500+ | $3,999 - $5,599 |10| Inference Stack | ROCm / llama.cpp | CUDA / TensorRT | Metal / llama.cpp |11+------------------------+-------------------+------------------+-------------------+
The NVIDIA DGX Spark utilizes an ARM-based Grace Blackwell configuration with higher memory bandwidth, granting it higher raw token-generation speeds on dense models. However, the Ryzen AI Halo runs on standard x86_64 architecture. This makes it substantially easier to integrate into existing Linux x86 developer toolchains, containerized CI/CD pipelines, and Windows environments without cross-compiling ARM binaries. For teams requiring native x86 compatibility alongside a 128GB GPU footprint, the Halo offers lower software friction.
Apple silicon (such as the M2 or M4 Ultra) provides significantly higher memory bandwidth (800 GB/s), translating to faster tokens per second on 70B models. However, the Mac Studio operates exclusively within macOS and the Apple Metal software stack. The Ryzen AI Halo runs native Linux (including AMD's dedicated Debian-derived AI platform) with direct ROCm and open-source driver support. For engineers deploying services targeting cloud Linux environments, developing on native Linux x86 hardware avoids platform-specific quirks that can appear when using Metal or Darwin runtimes.
For practitioners looking for the best AI chip for local deployment where raw memory capacity, native x86 operating system flexibility, and high-speed local networking are mandatory, the Ryzen AI Halo represents a cost-effective, self-contained solution.
The top models this device can run at 4-bit, ranked by fit and speed.
| Model | Grade | Speed | VRAM |
|---|---|---|---|
| Holo4-35B-A3BHcompany | AA | 88.1 tok/s | 2.3 GB |
| Qwen3-30B-A3BAlibaba | AA | 38.3 tok/s | 5.4 GB |
| LFM2.5-8B-A1BLiquid AI | AA | 70.9 tok/s | 2.9 GB |
| Apertus 8BEPFL, ETH Zurich, CSCS | AA | 38.1 tok/s | 5.4 GB |
| Llama 3 8B InstructMeta | AA | 36.4 tok/s | 5.7 GB |

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.