Advertising disclosure: we earn commissions when you shop through the links below.
The ASUS ProArt GR1X is a compact desktop PC built on the NVIDIA RTX Spark Superchip, which pairs a 6144-core NVIDIA Blackwell RTX GPU with a 20-core NVIDIA Grace CPU. It is aimed at running local AI agents, model inference and creative work on Windows, with up to 128GB of unified memory and support for up to 120B-parameter LLMs. ASUS rates it at up to 1 petaflop of FP4 AI performance and a 140W TDP in a 150 x 150 x 51 mm chassis that can drive four 4K displays. Preorders opened on October 7, 2026; ASUS has not announced a price.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
The ProArt GR1X represents a distinct architectural shift in compact desktop hardware, engineered specifically for local inference, multi-agent frameworks, and developer workflows. Built around NVIDIA's RTX Spark platform (the N1X superchip), the system integrates a 6,144-core Blackwell RTX GPU with a 20-core ARM-based NVIDIA Grace CPU over a high-bandwidth NVLink-C2C interconnect. Packed into a 150 x 150 x 51 mm chassis displacing just 1.1 liters, it delivers workstation-tier compute in an enclosure smaller than most external GPU enclosures.
For developers and machine learning practitioners, the GR1X addresses the primary bottleneck of desktop AI: memory capacity. While standard consumer GPUs top out at 24GB or 32GB of VRAM, the GR1X provides up to 128GB of high-speed unified LPDDR5X memory directly accessible to both the CPU and GPU. Operating natively on Windows 11 with full CUDA and TensorRT support, the system bridges the gap between power-hungry multi-GPU desktop towers and memory-constrained small-form-factor PCs.
Positioned as an always-on workstation for agentic workloads and local model execution, the GR1X competes directly with high-end Apple Silicon desktops and bespoke multi-GPU workstations. For engineers evaluating custom setups or emerging unknown hardware for AI development, the ProArt GR1X sets a compelling benchmark for power-to-footprint efficiency, combining 1 petaflop of FP4 AI compute with a restrained 140W TDP.
The compute engine inside the ProArt GR1X is the NVIDIA RTX Spark superchip, which pairs NVIDIA's Blackwell architecture with Grace server silicon. By routing data between the 20-core Grace CPU and the 6,144-core Blackwell GPU through NVLink-C2C, the system eliminates the standard PCIe bus penalty, granting the GPU zero-copy access to the full pool of unified memory.
Key technical specifications include:
ProArt GR1X AI inference performance benefits heavily from Blackwell's dedicated FP4 Tensor Cores. With up to 1 petaflop of low-precision compute, matrix multiplication operations execute with significantly lower latency and power overhead compared to prior-generation Ada Lovelace silicon. Because memory bandwidth dictates autoregressive token generation speeds, the unified LPDDR5X bus provides the sustained throughput needed to stream weights without the severe thermal throttling typical of mobile chips.
Thermal management is handled by an internal dual-fan assembly featuring 218 ultra-thin blades. The cooling system maintains a 140W envelope, allowing sustained, unthrottled execution during multi-hour batch runs or continuous agent polling. For practitioners looking for the best hardware for running AI models locally without dedicating a 1000W circuit to an inference rig, the efficiency profile here is an obvious draw.
Having an effective 128GB GPU for AI workloads changes what is possible on a desktop machine. Traditional consumer graphics cards force practitioners to aggressively quantize weights or offload layers to system RAM, which ruins generation speeds. The ProArt GR1X VRAM for large language models allows you to load dense models and large mixtures of experts fully onto the accelerator.
With 128GB of unified memory, system allocation overhead leaves approximately 110GB to 118GB for weights, context window buffers, and runtime activations:
Because Blackwell hardware natively accelerates FP4 execution, the sweet spot for maximum ProArt GR1X tokens per second is 4-bit weight precision. Using quantized formats like AWQ, NVFP4, or GGUF INT4, users can expect responsive generation across varied model classes:
Beyond language models, the 128GB pool allows concurrent execution of multimodal vision pipelines (like Qwen2-VL or LLaVA), audio models like Whisper-large-v3, and image generation architectures like Flux.1 Dev without needing to clear models from memory between pipeline steps.
The ProArt GR1X is built for environments where cloud API dependencies create latency, security, or recurring cost problems. It acts as a dedicated node for local machine learning workflows.
When selecting the best AI chip for local deployment, practitioners typically evaluate three hardware classes: Apple Silicon workstations, multi-GPU PC rigs, and compact superchips like the RTX Spark.
The Mac Studio has long been the primary choice for practitioners running large models locally due to its unified memory options (up to 192GB). However, the ProArt GR1X holds a crucial software advantage: native NVIDIA architecture.
The Apple ecosystem relies on MLX or Metal Performance Shaders (MPS), which frequently lag behind CUDA in kernel optimization, library updates, and flash attention support. The GR1X runs native CUDA, TensorRT, and vLLM without porting or framework workarounds. Additionally, Blackwell's native FP4 Tensor Cores provide superior low-precision throughput compared to Apple's Neural Engine and GPU compute. For teams strictly committed to the NVIDIA software ecosystem, the GR1X offers the unified memory benefits of a Mac Studio without leaving the CUDA platform.
A traditional desktop workstation containing two 24GB or 32GB GPUs provides raw FP16/FP8 compute, but splits memory across physical PCIe slots. Running a 70B or 120B model across two cards requires pipeline or tensor parallelism, introducing bus bottlenecks and physical footprint challenges.
A dual-GPU rig easily demands 800W to 1200W of power, requires a massive tower case, and generates severe heat and noise. The ProArt GR1X delivers a unified 128GB pool in a 1.1-liter case operating at a modest 140W TDP. While high-end discrete GPUs may yield higher raw bandwidth for smaller models, the GR1X is far more practical for running massive models locally in quiet office or home lab environments.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.