Advertising disclosure: we earn commissions when you shop through the links below.
Microsoft built the Surface Laptop Ultra with NVIDIA as its fastest Surface Laptop, aimed at developers, creators and people running AI models locally. It combines an NVIDIA Blackwell RTX Spark GPU with a 20-core Grace CPU and up to 128 GB of unified memory with full CUDA support. Microsoft says it reaches 1 petaflop of AI compute and can run models up to 120 billion parameters on the device. The 15-inch mini-LED PixelSense Ultra touchscreen is rated at 2,000 nits peak HDR brightness and 262 ppi. It ships in Platinum and Nightfall finishes, is listed for a 2026-09-27 release, and pricing was not announced.
Sized for production serving of 70B–200B class models at full or lightly-quantized precision. Overkill for a homelab; right call when the workload pays for itself in token volume.
Generated from this product’s spec sheet. Editor reviews refine it over time.
Microsoft's Surface Laptop Ultra is the flagship of the NVIDIA RTX Spark platform, and it is the first laptop to pair a 20-core Grace CPU with a Blackwell-generation GPU under full CUDA support, backed by up to 128 GB of unified memory. Announced at Computex 2026 with a listed release date of 2026-09-27 and no pricing disclosed, it sits in a tier above every Copilot+ PC on the market today. This is not a thin-and-light with an NPU bolted on. It is a desktop-class AI dev box in a 15-inch aluminum chassis, positioned against the MacBook Pro with M-series Max silicon and against NVIDIA's own DGX Spark.
For anyone searching for the best hardware for running AI models locally, the pitch is straightforward. Microsoft claims 1 petaflop of AI compute and on-device model support up to 120 billion parameters. That combination has, until now, meant either a desktop workstation or a cloud instance. The fact that CUDA runs natively on this machine matters more than the raw numbers: the entire NVIDIA software stack, PyTorch CUDA kernels, TensorRT, llama.cpp CUDA backend, cuDNN, and the ecosystem of quantized kernels built on top of them, is available without emulation layers or ROCm roulette.
At 128 GB of unified memory, this is a genuinely unusual laptop. Whether it is the right purchase for your workload depends on whether your bottleneck is capacity or bandwidth.
The specs that determine AI inference throughput on the Surface Laptop Ultra are memory capacity, memory bandwidth, and GPU core count, in roughly that order of importance for single-user workloads.
The platform is closely related to the GB10 "Grace Blackwell" superchip in the DGX Spark, reworked for Windows on Arm with an NPU for Copilot+ features. Memory bandwidth is the number that matters most for token generation, since autoregressive decoding reads the entire active weight set from memory for every token produced. NVIDIA has not published unified memory bandwidth for the RTX Spark N1X, so treat published tok/s claims with caution until independent testing lands. What is known is that this is an LPDDR5X-class shared pool in a 15-inch laptop thermal envelope, not an HBM part. Sustained clocks will be well below what a desktop or server GPU holds.
That framing matters: 128 GB of VRAM gives you capacity, not datacenter throughput. You will fit large models, and you will run them at conversational speeds rather than batch-serving speeds.
With 128 GB of unified memory, expect roughly 110 GB to 120 GB usable for weights after the OS and KV cache. That is enough for essentially every open-weight model at the 120B class and below at 4-bit quantization, and for most 70B-class models at 8-bit.
| Model | Quantization | Approx. weight size | Fits in 128 GB |
|---|---|---|---|
| Llama 3.1 8B | FP16 | ~16 GB | Yes, with room for long context |
| Qwen 2.5 32B / Qwen3 32B | Q4_K_M | ~20 GB | Yes, easily |
| Mistral Large 123B | Q4_K_M | ~70 GB | Yes |
| Llama 3.3 70B / Llama 3.1 70B | Q4_K_M | ~40 GB | Yes, up to very long context |
| Llama 3.1 70B | Q8_0 | ~70 GB | Yes |
| Llama 3.1 70B | FP16 | ~140 GB | No |
| Mixtral 8x22B | Q4_K_M | ~80 GB | Yes |
| gpt-oss-120b | MXFP4 | ~63 GB | Yes, comfortably |
| DeepSeek-R1 (671B) | Q4_K_M | ~400 GB | No |
The headline "120 billion parameters" case is almost certainly the mixture-of-experts shape: gpt-oss-120b or a comparable MoE checkpoint that activates a small fraction of its parameters per token. That is the ideal workload for this machine, because MoE decoding reads only the active experts, so throughput stays high even though the full model sits in memory. A dense 120B model at 4-bit would technically load, but decode speed would drop into the low single digits.
Expected tokens per second for llama.cpp-class CUDA inference. Independent benchmarks do not exist yet and early hands-on sessions were supervised, so treat these as bandwidth-class estimates:
Quantization sweet spot: Q4_K_M for anything above 30B. The quality loss against FP16 is small for most tasks and the memory savings are what make the larger models fit at all. For 8B-32B models, Q8_0 is worth the extra gigabytes if you have them. Keep the KV cache in FP16 for short contexts and switch to Q8 KV cache for 64K-plus contexts, since Llama 3.1 70B at FP16 KV cache consumes roughly 0.3 MB per token across its 80 layers, which is over 40 GB at 128K tokens.
Multimodal and long-context work is viable. Vision-language models in the 7B-90B range fit alongside their vision towers, Whisper handles transcription on GPU, and diffusion models (SDXL, Flux) run on the same CUDA cores. Agentic workloads with tool calling and multi-turn context are a good fit, provided you budget KV cache headroom.
Local coding agents and chat: this is the strongest argument for the machine. A 70B or 120B MoE model running at usable speed with CUDA acceleration is a different experience from a 7B model on integrated graphics. If you run agentic coding workflows locally and want them off the cloud, this is one of the few laptops that can do it.
Application developers: CUDA support means PyTorch, TensorRT, and the standard toolchain are available. The caveat is Windows on Arm: verify ARM64 builds exist for your dependencies, and expect some Python wheels to need building from source.
Fine-tuning: QLoRA and LoRA on 7B-32B models is realistic. Full-parameter fine-tuning is not, and pretraining is out of scope entirely. This is an inference-first device.
Creators: Adobe Premiere Pro and Unreal Engine 5 both run on this hardware, and the 2,000-nit mini-LED panel is a genuine asset for color and HDR work alongside the AI workloads.
Teams running inference servers: this is not that device. It is a single-user workstation laptop. If you need concurrent request handling, look at a DGX Spark or a proper GPU server.
Versus the NVIDIA DGX Spark: same memory capacity class, same Grace Blackwell lineage, very different form factors. The DGX Spark is a fixed desktop box with a validated Linux software stack and predictable sustained clocks. The Surface adds a 15-inch 2,000-nit display, a battery, and portability at the cost of thermal headroom and Windows on Arm friction. Choose the Surface if the machine travels with you; choose the DGX Spark if it sits on a desk and you want the least software friction.
Versus a MacBook Pro with M-series Max and 128 GB unified memory: Apple's advantage is software maturity for local inference (MLX, llama.cpp Metal) and efficiency per watt. Its disadvantage is that CUDA-only kernels, TensorRT, and most of the NVIDIA-specific quantization tooling simply do not exist there. If your stack is CUDA-shaped, the Surface is the only laptop-class option at this memory tier today.
Versus a Ryzen AI Max+ 395 machine with 128 GB: usually cheaper, with comparable unified memory capacity, but the software story runs through ROCm rather than CUDA. For practitioners who want the least friction for local LLM deployment, CUDA support is the deciding factor.
Pricing was not announced, and that will determine whether this lands as a value proposition or a halo product. On capability alone, the Surface Laptop Ultra is the first Windows laptop that credibly answers the question of what to buy for local AI agents in 2026.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.

Rent this class of GPU by the hour before you buy. Live prices from RunPod and Vast.ai.

Find the break-even between buying this hardware and paying for a cloud API.
Mac vs NVIDIA for local inference, if you are still choosing a platform.