Advertising disclosure: we earn commissions when you shop through the links below.
Thunderobot's STATION Series AI Workstation is a 2-liter mini PC built around the AMD Ryzen AI Max+ 395 with Radeon 8060S graphics and a 50 TOPS XDNA 2 NPU, for 126 TOPS of total AI performance. It ships with 128 GB LPDDR5X-8000 memory, of which up to 96 GB can be assigned as VRAM, plus a 1 TB PCIe Gen4 SSD and two spare M.2 slots. Cooling uses four heat pipes and three fans in an aluminum alloy frame to sustain a 132 W TDP, and networking includes dual 10GbE, Wi-Fi 7 and Bluetooth 5.4. It is sold in China at CNY 26,999, discounted to CNY 25,499, with pre-orders open and sales starting October 9; a Ryzen AI Max+ Pro 495 version is planned.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
Thunderobot's STATION Series AI Workstation is a 2-liter mini PC built around AMD's Ryzen AI Max+ 395, the "Strix Halo" APU that pairs 16 Zen 5 cores with a 40-CU RDNA 3.5 integrated GPU and a 50 TOPS XDNA 2 NPU. What makes it interesting for AI work is not the compute, it is the memory: 128 GB of LPDDR5X-8000 on a 256-bit bus, with up to 96 GB assignable as VRAM to the Radeon 8060S. That single number puts it in a class of machine that, until this generation, effectively did not exist below workstation or data center pricing.
At $4,022 (CNY 26,999, currently discounted to CNY 25,499 in China), this is a prosumer and small-team inference box, not a consumer toy and not a rack-mounted server. It competes directly with other Strix Halo mini PCs such as the GMKtec EVO-X3 and MSI PRO MAX EDGE AI+, and indirectly with Apple's unified-memory Mac Studio. It is available for pre-order, with sales starting October 9; a Ryzen AI Max+ Pro 495 variant is planned.
The pitch is straightforward: 96 GB of addressable VRAM in a chassis the size of a large book, with dual 10GbE and a 132 W sustained TDP. For anyone running local LLMs, that memory ceiling is the whole story.
The specs that matter for inference, in order of how much they constrain your workload:
The distinction between compute and bandwidth matters here. 126 TOPS of INT8 throughput is respectable for a mini PC, but prompt processing (prefill) is compute-bound while token generation (decode) is bandwidth-bound. At ~256 GB/s, you should expect decode speeds in the range of 5-12 tokens/second for dense 70B-class models at 4-bit, and considerably faster for 8B-14B models. The NPU's 50 TOPS is useful for smaller always-on workloads and Windows Copilot+ features, but most local LLM runtimes still target the GPU, so plan around the 8060S.
Power efficiency is a genuine strength. A 132 W envelope for 96 GB of usable VRAM is roughly an order of magnitude below a multi-GPU workstation doing the same job.
With 96 GB of VRAM, this machine runs a broader range of models than any single consumer GPU. Sizing at 4-bit (Q4_K_M), which is the practical default for local deployment:
The quality-to-speed sweet spot is 4-bit for anything above 30B, and 8-bit for 7B-14B models where you have the memory to spare. Going to 6-bit for a 70B model is possible but leaves little room for KV cache.
Multimodal models (Qwen2-VL, Llama 3.2 Vision, LLaVA) run fine within the memory budget. Long-context work is where the unified memory pool pays off: a 128K-context run on a 70B model needs tens of gigabytes of KV cache, which is a non-starter on a 24 GB or 48 GB GPU but manageable here.
Local LLM hobbyists and researchers who want to run 70B-class models at home without a multi-GPU tower. This is the primary buyer.
AI engineers building agentic workflows where multiple models (a planner, a coder, a summarizer) need to be resident simultaneously. 96 GB lets you keep several models loaded at once instead of swapping.
Small teams running inference servers for internal tools. Dual 10GbE is the tell: this is meant to sit on a network and serve tokens to other machines, not just one desk.
Edge and on-prem deployment where data cannot leave the building. The 2-liter footprint and 132 W draw make it deployable in offices, clinics, or labs without a server room.
Not for training. Fine-tuning small models (LoRA on 7B-13B) is feasible. Full-parameter training of anything substantial is not what this hardware is for. Treat it as an inference and light-adaptation machine.
vs. GMKtec EVO-X3 / MSI PRO MAX EDGE AI+: These are the closest direct competitors, also Strix Halo with 96 GB. Differences come down to cooling, port selection, and price. Thunderobot's dual 10GbE and 132 W sustained TDP are the differentiators if you need network serving or sustained load without throttling. Pick the Thunderobot when networking and thermal headroom matter; pick a cheaper Strix Halo box if you only need one machine on a desk.
vs. Apple Mac Studio (M4 Max / M3 Ultra): Apple's unified memory architecture is the other way to get large VRAM without a discrete GPU. The Mac Studio offers higher memory bandwidth on the Ultra tier and a more mature software stack for MLX. The Thunderobot runs Windows and CUDA-adjacent tooling (ROCm, Vulkan, DirectML) more naturally, has dual 10GbE, and is upgradeable on storage via the spare M.2 slots. Choose Apple for bandwidth-bound decode speed and macOS tooling; choose the STATION for Windows/ROCm workflows and network-centric deployment.
vs. a used dual-3090 or single-4090 build: Two 3090s give 48 GB and much higher bandwidth, but at 700+ W, in a full tower, with no CPU/NPU integration and no warranty. The STATION trades raw throughput for density, power, and a single warranty. If you need maximum tokens/second and have the space, GPUs win. If you need 96 GB in 2 liters, this wins.
The Ryzen AI Max+ Pro 495 variant, when it ships, will likely widen the gap on the compute side without changing the memory story. For now, the STATION Series AI Workstation is one of the few ways to get 96 GB of GPU-addressable VRAM in a package you can carry under one arm.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.