Advertising disclosure: we earn commissions when you shop through the links below.
Framework Desktop is a 4.5L mini workstation from Framework built around the AMD Ryzen AI Max+ PRO 495 with 192GB of soldered LPDDR5X-8533 memory on a 256-bit bus and 273 GB/s of memory bandwidth. It pairs 16 Zen 5 CPU cores with a 40-core Radeon 8065S GPU and an NPU rated up to 55 TOPS, and Framework says the 192GB configuration can run 300B-class open-weight models at 4-bit quantization. Pre-orders opened on 2026-09-30 at $6,799 for the DIY Edition and $7,449 for the 2TB pre-built, with shipments starting in November. The Mini-ITX system takes two PCIe 4.0 M.2 drives up to 8TB each and includes USB4, 5Gbit Ethernet and Wi-Fi 7.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
The Framework Desktop with the AMD Ryzen AI Max+ PRO 495 and 192GB of unified LPDDR5X answers one question: how large a model can you run on a desk without a data center? Framework's answer is 192GB of soldered memory on a 256-bit bus delivering 273 GB/s of bandwidth, paired with 16 Zen 5 CPU cores, a 40-core Radeon 8065S GPU, and an NPU rated at up to 55 TOPS. The DIY Edition is $6,799, a pre-built with a 2TB NVMe drive and Fedora pre-installed is $7,449, and pre-orders opened on 2026-09-30 with shipments starting in November.
This is prosumer and small-team hardware, not data center gear. The entire system draws 120W sustained (140W boost) from a 400W FlexATX supply, weighs 3.1kg, measures 96.8 x 205.5 x 226.1mm, and runs Linux. What it buys you is capacity. An RTX 5090 has 32GB of VRAM and roughly 6.5x the bandwidth. This machine has 192GB of it. For dense models above 70B parameters at aggressive quantization, and for mixture-of-experts models that are cheap to compute but enormous to store, capacity is the binding constraint, not raw FLOPS.
Framework says the 192GB configuration runs 300B-class open-weight models at 4-bit quantization, and specifically cites DeepSeek-V4.1-Flash at Q2 and MiMo V2.6 Flash RL at MXFP4 as working examples. Treat those as vendor claims until independent benchmarks land, but the underlying math holds: 192GB holds weights that no single consumer GPU, and few dual-GPU workstations, can touch.
Decode throughput scales roughly with bandwidth divided by bytes read per token. At 273 GB/s and a realistic 55 to 70 percent efficiency, a 70B model at 4-bit (roughly 40GB of weights) lands in the 4 to 7 tokens/second range. A 32B model at 4-bit reaches the low teens. Small 8B models are compute-limited rather than bandwidth-limited, so expect tens of tokens per second, capped by the Radeon 8065S rather than by memory.
Prefill is the opposite problem. Prompt processing is compute-bound, and 40 RDNA 3.5 cores will chew through long contexts noticeably slower than a discrete GPU. Long-context work is viable because you have memory for the KV cache, but time-to-first-token runs into seconds on large prompts.
For context: NVIDIA's DGX Spark matches the 273 GB/s but tops out at 128GB. Apple's M3 Ultra Mac Studio offers 819 GB/s and up to 512GB at a much higher price point, without x86 Linux. An RTX 5090 delivers roughly 1,792 GB/s but only 32GB.
Q4_K_M (or MXFP4 on models trained for it) is the right default. Quality degradation is small and weight size drops enough that context and KV cache still fit. Drop to Q3 or Q2 only when capacity forces it, and expect measurable loss on reasoning tasks. For 30B-class models, Q6_K is often worth the extra memory because there is room.
Connectivity supports the agent use case: dual USB4 Type-C, dual USB-A 3.2 Gen 2, HDMI 2.1, two DisplayPort 2.1 (UHBR10), 5Gbit Ethernet, Wi-Fi 7, a 3.5mm combo jack, and two front Expansion Card slots.
vs NVIDIA DGX Spark (128GB): Same 273 GB/s bandwidth, same unified-memory philosophy, same target workload. The Spark ships with CUDA and the NVIDIA software stack, which is a real advantage if you are already in that ecosystem. The Framework gives you 64GB more memory, x86 Linux, swappable storage, and no runtime lock-in. If your models fit in 128GB, the Spark is the easier buy. If they do not, this is the only option near the price.
vs Mac Studio M3 Ultra: Apple's machine offers 819 GB/s and up to 512GB of unified memory, and it will generate tokens several times faster on models that fit. It is also more expensive once you configure large memory, runs macOS, and shuts you out of ROCm and the wider x86 Linux tooling. Choose the Mac for raw speed on Apple-supported runtimes. Choose the Framework for Linux, x86 compatibility, and a lower entry price at the 192GB tier.
vs dual RTX 5090: Two 5090s cost less than this system and deliver several times the compute and bandwidth, but only 64GB of combined VRAM. They will crush 30B models and choke on 120B MoE. Different tools for different constraints.
Pick the Framework Desktop when your bottleneck is how much model fits, not how many tokens per second it emits.
The top models this device can run at 4-bit, ranked by fit and speed.
| Model | Grade | Speed | VRAM |
|---|---|---|---|
| Qwen3-30B-A3BAlibaba | AA | 40.8 tok/s | 5.4 GB |
| Holo4-35B-A3BHcompany | AA | 94.0 tok/s | 2.3 GB |
| LFM2.5-8B-A1BLiquid AI | AA | 75.6 tok/s | 2.9 GB |
| Llama 3 8B InstructMeta | AA | 38.8 tok/s | 5.7 GB |
| LensVLM-9BApple | AA | 36.5 tok/s | 6.0 GB |

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.