Advertising disclosure: we earn commissions when you shop through the links below.
Minisforum's MS-S1 MAX-P495 is an expandable mini workstation built around AMD's Ryzen AI Max+ PRO 495 with 16 Zen 5 cores and Radeon 8065S graphics. It carries up to 192GB of LPDDR5X unified memory, of which up to 160GB can be allocated as VRAM, and delivers up to 131 TOPS of total AI compute including a 55 TOPS NPU. Minisforum targets local LLM inference and creative work, and says a four-unit cluster can run DeepSeek-R1 671B at Q4_0 quantization. Pricing has not been announced and the listed release date is 2026-09-25.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
Minisforum's MS-S1 MAX-P495 is an expandable mini workstation built around AMD's Ryzen AI Max+ PRO 495 APU. It pairs 16 Zen 5 cores with Radeon 8065S graphics and a dedicated NPU, but the headline for AI workloads is memory: up to 192GB of LPDDR5X-8533 unified memory, with up to 160GB allocatable as VRAM. That puts it in a rare tier of local inference hardware — prosumer and small-team workstation, not data center, not consumer mini PC.
For anyone searching for the best hardware for running AI models locally, the MS-S1 MAX-P495 addresses the primary bottleneck: VRAM capacity. A single unit can hold models that would otherwise require multiple discrete GPUs or a cloud instance. And because it supports 2U rack deployment and multi-unit clustering, Minisforum claims a four-unit cluster can run DeepSeek-R1 671B at Q4_0 quantization (380GB total). That's a 671B parameter model on four desktop-class boxes — a configuration previously reserved for server racks.
The system is announced with a listed release date of 2026-09-25. Pricing has not been announced. It competes most directly with the Mac Studio M3 Ultra (192GB unified memory) and Nvidia's DGX Spark (128GB unified memory), but differentiates on x86 compatibility, rack support, and a 55 TOPS NPU. If you're evaluating the MS-S1 MAX-P495 for AI, the key question is whether its unified memory architecture and clustering capability fit your inference workload better than a discrete-GPU build or an ARM-based alternative.
The specs that matter for local LLM inference are memory capacity, memory bandwidth, and compute throughput. Here's how the MS-S1 MAX-P495 breaks down:
Memory bandwidth is not listed, but LPDDR5X-8533 on a wide bus typically delivers around 250–270 GB/s. That's lower than an RTX 4090 (1,008 GB/s) or Mac Studio M3 Ultra (800 GB/s). For token generation, bandwidth is the limiting factor. Expect the MS-S1 MAX-P495 to generate tokens slower than high-end discrete GPUs on models that fit in both, but to run models that simply won't fit on a 24GB or 48GB card. The trade-off is capacity over raw speed.
The 55 TOPS NPU can offload supported operations, but framework support for AMD's NPU in LLM inference is still maturing. For now, treat the NPU as a bonus for specific workloads rather than the primary inference engine.
This is where the MS-S1 MAX-P495 earns its place. With 160GB of allocatable VRAM, you can run:
For the MS-S1 MAX-P495, 4-bit quantization (Q4_K_M, Q4_0, AWQ) offers the best quality-to-speed tradeoff for models above 30B. At 4-bit, a 70B model uses ~40GB, leaving 120GB for context and other models. 8-bit is viable for models up to ~140B, but token generation will be slower due to higher memory traffic. For models under 30B, 8-bit or even 16-bit is practical.
Minisforum has not published MS-S1 MAX-P495 tokens per second benchmarks. Based on similar unified-memory architectures (AMD Strix Halo, Apple M-series), expect:
These are estimates. Actual performance depends on batch size, framework (llama.cpp, vLLM, ROCm), and whether the NPU is used. For interactive chat, 5–12 tokens/s on a 70B model is usable. For batch inference, throughput scales with batching.
Hobbyists and power users running local chatbots or experimenting with cutting-edge models. The 160GB VRAM ceiling lets you load a 70B model and keep it resident without quantization compromises. This is one of the best hardware options for local AI agents in 2026 if your agent needs to call multiple models simultaneously.
Developers building AI-powered applications who need a local inference server. The x86 architecture means standard Linux and Windows toolchains work. ROCm support for Radeon 8065S is improving but still less mature than CUDA. If your stack is PyTorch + ROCm or llama.cpp, this is viable.
Teams running inference servers that need to serve multiple users or models. The 2U rack support and clustering capability make it deployable in a small office or lab. Four units give you 640GB of total VRAM for 671B-class models.
Edge deployment where power and space are constrained. At 160W peak, this is far more efficient than a multi-GPU server. You can run it on a desk or in a rack without dedicated cooling.
Training vs. inference: This is an inference machine. Fine-tuning small models (7B–13B) with LoRA is possible, but full fine-tuning of large models is not practical. No NVLink, no high-bandwidth GPU interconnect. Buy it for inference and experimentation, not training.
MS-S1 MAX-P495 vs Mac Studio M3 Ultra: Both offer up to 192GB unified memory. The Mac Studio has higher memory bandwidth (~800 GB/s) and a more mature software stack for creative work. The MS-S1 MAX-P495 runs x86, supports 2U rack deployment, and has a 55 TOPS NPU. If you need CUDA-adjacent tooling or rack mounting, pick the Minisforum. If you need maximum token generation speed on models that fit in 192GB, the Mac Studio wins.
MS-S1 MAX-P495 vs Nvidia DGX Spark: DGX Spark offers 128GB unified memory and 1 PFLOP sparse FP4 compute, but it's ARM-based and has less memory. The MS-S1 MAX-P495 gives you 160GB VRAM and x86 compatibility. DGX Spark is more AI-optimized out of the box; the MS-S1 MAX-P495 is more flexible for general workstation workloads.
When to pick the MS-S1 MAX-P495: You need more than 128GB of VRAM, you want x86, you plan to cluster multiple units, or you need rack deployment. It's the best AI chip for local deployment when capacity matters more than raw speed.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.

Get a full budget-matched parts list for a local AI workstation.
Mac vs NVIDIA for local inference, if you are still choosing a platform.