Advertising disclosure: we earn commissions when you shop through the links below.
GMKtec's EVO-X5 Pro is a compact desktop AI PC built around AMD's Ryzen AI Max+ PRO 495 processor with up to 192 GB of unified LPDDR5X memory. It is aimed at running large language models locally, with up to 160 GB of that memory allocatable as graphics memory and support for 300-billion-parameter models at INT4 quantization. The chip pairs a 16-core Zen 5 CPU, Radeon 8065S integrated graphics with 40 compute units, and an XDNA 2 NPU rated at 55 TOPS. GMKtec announced it at IFA 2026 with a launch date of 28 September 2026; pricing has not been shared.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
GMKtec's EVO-X5 Pro is a compact desktop AI workstation built around AMD's Ryzen AI Max+ PRO 495, the "Gorgon Halo" SoC that pairs 16 Zen 5 cores with a 40-compute-unit Radeon 8065S iGPU and an XDNA 2 NPU. The headline number is memory: up to 192 GB of unified LPDDR5X, with as much as 160 GB allocatable to the GPU. That single figure is what separates this machine from nearly everything else in the mini PC class and puts it in the same conversation as NVIDIA's DGX Spark and Apple's high-memory Mac Studio.
GMKtec announced the EVO-X5 Pro at IFA 2026 in Berlin with a launch date of 28 September 2026. Pricing has not been shared, which is the one missing variable that determines whether this lands as a bargain or a curiosity. The positioning is explicit: a local platform for agentic workloads, where models are hosted on the machine, agents are orchestrated on the machine, and nothing needs to leave the desk.
Tier-wise, this is prosumer to small-team hardware. It is not a data center accelerator and it will not train a frontier model. What it does is hold a 300-billion-parameter model at INT4 in memory and serve it locally, which until recently required either a 512 GB Mac Studio or a multi-GPU server rack.
The specs that matter for inference, in order of importance:
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.
Generated from this product’s spec sheet. Editor reviews refine it over time.
Memory bandwidth is the number that determines token generation speed for autoregressive decoding. At roughly 273 GB/s, the EVO-X5 Pro sits in Apple M4 Pro territory and well below the M3 Ultra's 819 GB/s. That gap is the single biggest constraint on this machine: a 70B model at Q4 fits comfortably in 40 GB, but every token generated has to stream those weights through a much narrower pipe than a Mac Studio offers. Where the EVO-X5 Pro wins is capacity per dollar and x86 software compatibility, not raw bandwidth.
The 55 TOPS NPU is useful for the workloads that actually target it: Windows Studio Effects, some ONNX Runtime and DirectML paths, and lightweight always-on inference. Most LLM serving stacks (llama.cpp, vLLM, Ollama, LM Studio) run on the GPU via ROCm or Vulkan and will not touch the NPU. Treat the NPU as a bonus for edge vision and audio pipelines, not as the reason to buy.
This is where the 160 GB GPU allocation earns its keep. Practical model fit at common quantization levels:
Expected tokens/second for popular models, given the bandwidth ceiling (treat these as estimates, not benchmarks):
The sweet spot on this hardware is Q4_K_M or INT4 for anything above 30B. Q5 and Q6 give better output quality but cut your usable model ceiling and slow generation proportionally. For models at 70B and below, Q5_K_M is a reasonable trade if you value quality over speed. For 200B+ models, INT4 is the only option that fits at all, and you should expect single-digit token rates suited to batch or agentic work rather than interactive chat.
Multimodal models are viable. Vision-language models in the 7B-34B range (Qwen2.5-VL, Llama 3.2 Vision) fit easily, and diffusion models like SDXL and Flux.1 Dev run comfortably in the memory budget. Long-context work is limited mainly by KV cache growth: at 128K context on a 70B model, the cache alone can consume tens of gigabytes, so budget accordingly against your 160 GB ceiling.
Local chatbot and assistant hosting. If you want a private ChatGPT replacement that runs 70B-class models at usable speed without a subscription, this is a strong fit. The 160 GB pool means you can keep multiple models resident and swap between them without reloading from disk.
Developers building agentic applications. GMKtec's pitch is the agentic PC, and the hardware backs it: enough memory to hold a large reasoning model plus tool-calling models plus embeddings simultaneously, with DASH remote management for headless operation. Teams building local AI agents in 2026 should shortlist this against DGX Spark.
Small-team inference servers. With a dTPM 2.0 chip, RJ45 DASH, and up to 24 TB of NVMe storage, this is deployable as a departmental inference node. Throughput per box is modest, but the cost per concurrent agent session is likely far below cloud API pricing at sustained volume.
Edge and air-gapped deployment. Fully offline operation is the core selling point. Regulated environments, field sites, and anything where data cannot leave the building are the obvious markets.
Training: no. Fine-tuning a 7B model with LoRA is possible but slow. Full fine-tuning is out of scope. Buy this for inference.
vs. NVIDIA DGX Spark (GB10). DGX Spark offers 128 GB of unified memory and a CUDA software stack that is still the default for serious ML tooling. The EVO-X5 Pro counters with 192 GB total and 160 GB GPU-allocatable, meaning it can hold larger models at INT4 than Spark can. GMKtec claims 66% faster CPU orchestration and 40% lower cost per completed AI workflow versus DGX Spark, though those are vendor figures. If your stack depends on CUDA, Spark wins on ecosystem. If you need maximum model capacity and x86 flexibility, the EVO-X5 Pro has the edge.
vs. Apple Mac Studio (M3 Ultra, 512 GB). The Mac Studio has roughly 3x the memory bandwidth and a much higher memory ceiling, and macOS with MLX is a mature local inference platform. It costs substantially more. The EVO-X5 Pro is the better pick when you want x86 software compatibility, DASH management, and eGPU expansion via USB4 V2, and when 160 GB is enough.
vs. other Ryzen AI Max+ mini PCs. Framework Desktop and HP's Z2 Mini G1a use the same silicon family. GMKtec differentiates on the 192 GB configuration, triple M.2 storage, and the management features. Compare carefully on price once GMKtec publishes it, since the underlying compute is largely identical across vendors.