Advertising disclosure: we earn commissions when you shop through the links below.
GMKtec's EVO-X3 is a compact AI workstation mini PC built around the AMD Ryzen AI Max+ 395 processor with Radeon 8060S integrated graphics and an XDNA 2 NPU. It ships with 128GB of onboard LPDDR5X-8000 memory, up to 96GB of which can be allocated as VRAM, and GMKtec lists support for LLMs up to 235B parameters. Total AI performance is stated at 126 TOPS, with 50 TOPS from the NPU. Pricing starts at $3,799 with a 2TB or 4TB SSD, and it runs Windows 11 Pro, Ubuntu or Linux.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
The EVO-X3 is GMKtec's second-generation Strix Halo mini workstation, built around AMD's Ryzen AI Max+ 395. It replaces the EVO-X2 with a redesigned full-metal CNC chassis, a beefier cooling system, and 128GB of LPDDR5X-8000 memory as standard. At $3,799.99 with a 2TB or 4TB NVMe drive, it sits firmly in the prosumer tier — the same bracket as Apple's Mac Studio M3 Ultra and high-end discrete-GPU workstations like a dual RTX 4090 build, but in a 353 x 186 x 41 mm, 2.3 kg enclosure.
What makes the EVO-X3 interesting for AI work isn't raw compute — it's the memory architecture. The Ryzen AI Max+ 395 uses a unified memory pool, so up to 96GB of the 128GB LPDDR5X can be carved out as VRAM for the Radeon 8060S iGPU. That's the headline number: 96GB of GPU-addressable memory in a mini PC, which is enough to load quantized 200B+ class models that simply won't fit on a single 24GB or 48GB consumer GPU. GMKtec lists support for LLMs up to 235B parameters.
For practitioners weighing the EVO-X3 for AI, the tradeoff is clear. You get enormous memory capacity and quiet operation at 140W peak, but you pay a premium and accept integrated-graphics-tier compute throughput. This is a capacity play, not a speed play.
The specs that matter for local inference:
The 50 TOPS NPU is useful for offloading specific workloads — Windows Studio effects, small on-device models via ONNX Runtime or Ryzen AI toolchains — but most LLM inference on Strix Halo runs on the Radeon 8060S through ROCm or Vulkan backends. The NPU isn't the LLM workhorse here; the iGPU and unified memory are.
Memory bandwidth is the constraint to understand. LPDDR5X-8000 delivers roughly 256 GB/s of theoretical bandwidth shared between CPU and GPU. That's excellent for an integrated part and comparable to an M4 Pro, but it's a fraction of what a discrete RTX 4090 (1 TB/s) or an H100 (3.35 TB/s) provides. The practical consequence: prompt processing (prefill) is fast, but token generation (decode) will be bandwidth-limited. Expect the EVO-X3 to be competitive on time-to-first-token and slower on sustained tokens/sec than discrete-GPU alternatives.
The 140W peak mode is worth enabling for sustained inference. The chassis handles it quietly — Tom's Hardware and TechRadar both flagged the cooling as a strength, with the machine staying quiet under load.
This is where the EVO-X3 justifies its price. With 96GB of VRAM, the model compatibility envelope is far wider than any single consumer GPU.
Comfortably fits (fast, high quality):
Fits with careful quantization:
Sweet spot: For most practitioners, Q4_K_M or Q5_K_M quantization of 70B–123B models offers the best quality-to-speed tradeoff. A 70B model at Q4_K_M (~40GB) leaves plenty of VRAM headroom for long context windows — 32K to 128K tokens is realistic without paging.
Multimodal and long-context: The unified memory pool makes the EVO-X3 unusually capable for vision-language models (Qwen2-VL, Llama 3.2 Vision, InternVL) and long-context work. Loading a 70B model plus a large KV cache for 128K context is feasible in a way it isn't on a 24GB card.
Expected tokens/sec: Community benchmarks on Strix Halo with 128GB typically land in the 5–10 tok/s range for 70B-class models at Q4, and 15–25 tok/s for 30B-class models. Smaller 7B–14B models run at 40+ tok/s. These are estimates — actual throughput depends on your runtime (llama.cpp, Ollama, LM Studio, vLLM), quantization, and context length.
Developers building agentic workflows. The 96GB VRAM pool lets you run a capable 70B model as a planner alongside smaller specialized models, all resident in memory simultaneously. For "best hardware for local AI agents 2026" searches, the EVO-X3's capacity is its differentiator — you're not juggling model loads.
Researchers and hobbyists running large local LLMs. If you want to run DeepSeek-R1-70B or Qwen 2.5 72B at high quality without a multi-GPU rig, this is one of the few single-box options under $4K.
Teams running inference servers. The 140W ceiling and quiet operation make it viable in an office or lab. OCuLink gives you a path to add a discrete GPU later if you need more throughput.
Edge deployment. Full metal chassis, 2.3 kg, Windows 11 Pro or Ubuntu. It's portable enough for field deployments where a rack isn't available.
Not for training. The 256 GB/s bandwidth and integrated GPU make this an inference machine. Fine-tuning a 7B model with LoRA is possible; anything larger is impractical.
vs. Apple Mac Studio M3 Ultra (128GB/256GB): The Mac Studio offers higher memory bandwidth (up to 819 GB/s on M3 Ultra) and a more mature software stack for MLX-based inference. It typically delivers faster token generation. The EVO-X3 counters with x86 compatibility, CUDA-alternative ROCm tooling, OCuLink expansion, and native Windows/Linux support. If your stack is MLX or you prioritize tok/s, the Mac wins. If you need x86 or plan to add an eGPU, the EVO-X3 is the pick.
vs. dual RTX 4090 build (~$4,500+): Two 4090s give you 48GB VRAM and dramatically higher compute throughput — 5–10x faster token generation on models that fit. But 48GB caps you at ~70B Q4, and you're dealing with 900W+ of heat and noise. The EVO-X3 runs 128B+ models in a 2.3 kg box at 140W. Different tools for different jobs: the 4090 rig is for speed on mid-size models, the EVO-X3 is for capacity on large ones.
vs. GMKtec EVO-X2: Same Ryzen AI Max+ 395 silicon, but the X3 adds the new chassis, standard 128GB LPDDR5X-8000, and improved cooling. The X2 launched at roughly half the price — the X3's premium reflects the 2026 component shortage as much as the redesign.
Pick the EVO-X3 if your workload is memory-bound, not compute-bound — large quantized models, long context, multimodal, or multi-model agent stacks. Skip it if you need maximum tokens/sec on models that already fit in 24–48GB.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.

Get a full budget-matched parts list for a local AI workstation.
Mac vs NVIDIA for local inference, if you are still choosing a platform.