Qwen3.8-Max is Alibaba's largest Qwen model, a sparse mixture-of-experts system with 2.4 trillion total parameters and about 95 billion active per token. It supports a 1M-token context window and native visual understanding, and is aimed at coding, long-horizon autonomous tasks and professional work such as legal, financial and design tasks. It is served through Alibaba Cloud Model Studio and QwenCloud at $2 per million input tokens and $6 per million output tokens. Alibaba says the open weights will be released the week after launch.
A workable 2.4B-parameter MoE language model from Alibaba. Pulls ahead on graduate-level reasoning (GPQA) (93/100), so reach for it when that's the dimension that matters.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 795.6 GB | Low | |
| Q4_K_MRecommended | 815.6 GB | Good | |
| Q5_K_M | 825.1 GB | Very Good | |
| Q6_K | 836.5 GB | Excellent | |
| Q8_0 | 860.2 GB | Near Perfect | |
| FP16 | 950.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| ASUS ExpertCenter Pro ET900N G3ASUS | DD | 7.0 tok/s | 815.6 GB |
| Dell Pro Max with GB300Dell | DD | 7.0 tok/s | 815.6 GB |
| Gigabyte W775-V10-L01Gigabyte | DD | 7.0 tok/s | 815.6 GB |
| HP ZGX Fury AI StationHP | DD | 7.0 tok/s | 815.6 GB |
| MSI XpertStation WS300MSI | DD | 7.0 tok/s | 815.6 GB |
Energy cost on MSI XpertStation WS300 (~7.0 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Qwen3.8-Max on MSI XpertStation WS300 · ~7.0 tok/s · 1400W | $6.66 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 816 GB VRAM, refreshed hourly.
No current rental listing covers this model’s VRAM requirement on the providers we track.
Qwen3.8-Max is Alibaba's flagship Qwen model, and it is the largest the team has shipped: a sparse mixture-of-experts system with 2.4 trillion total parameters and roughly 95 billion active per token. It is multimodal (text, images, video, and documents in, text out), carries a 1M-token context window, and is positioned against frontier closed models rather than the small local models Qwen is best known for. Alibaba serves it through Alibaba Cloud Model Studio and QwenCloud at $2 per million input tokens and $6 per million output tokens. Open weights are promised for the week after launch, with no licence specified at the time of writing.
The practical headline for anyone building on this: Qwen3.8-Max is a datacenter-class model. It leads Alibaba's published benchmark table on PaperBench (93.0) and IFBench (82.8), and it is competitive on GPQA Diamond (92.6). It also trails Claude Fable 5 on SWE-bench Pro by 12.3 points (67.7 vs 80.0) and on HLE (43.6 vs 53.3). Several of those rows are Alibaba's own internal benchmarks (QwenSWEBench, QwenQoderBench, CoWorkBench, SkillsBench), so treat them as directional rather than independently verified.
If you landed here because you want to run Qwen locally, read the Running Locally section before you get excited. The 2.4T figure is not a typo, and it changes everything about hardware requirements.
Qwen3.8-Max uses a sparse mixture-of-experts architecture. Total parameter count is 2.4T; active parameters per token are about 95B. That ratio (roughly 4% active) is what makes the model servable at all. A dense 2.4T model would be prohibitively slow and expensive; a sparse MoE routes each token through a small subset of experts, so compute scales with active parameters while capacity scales with total.
For inference, the number that governs speed is 95B active. That is still large by local standards: it is roughly the compute footprint of a dense ~95B model per token, well beyond any single consumer GPU. VRAM, however, is driven by total parameters, not active ones, because all experts must be resident (or paged) to route correctly. At Q4_K_M, 2.4T parameters land around 1.2TB of weights. At FP8, closer to 2.4TB. No consumer configuration holds this.
The 1M-token context window is the other defining spec. It enables whole-repository code analysis, long document sets (contracts, filings, design specs), and multi-hour agent traces without chunking. Alibaba lists 991K maximum input, 131K maximum output, and 262K maximum reasoning on the QwenCloud model page.
Listed capabilities: chat, code, vision, reasoning, function-calling.
You cannot run the full Qwen3.8-Max on consumer hardware. This is not a "needs a 4090" situation; it is a "needs a rack" situation.
VRAM requirements by precision (weights only, before KV cache and activations):
For reference, an RTX 4090 has 24GB, an M4 Max tops out at 128GB of unified memory, and a maxed Mac Studio M3 Ultra reaches 512GB. None of these fit even a Q2 build of the full model. Multi-GPU H100 or H200 nodes with 1.5TB+ of aggregate VRAM are the realistic floor.
Tokens per second depends almost entirely on the serving stack and interconnect. On an 8x H200 node with FP8 and expert parallelism, expect low double-digit to low triple-digit tok/s for single-stream generation, higher under batched serving. On anything less, you are paging experts off disk, and throughput collapses.
Best quantization for Qwen3.8-Max: FP8 is the practical production choice. Q4_K_M is only relevant if you are running the open weights on a very large unified-memory box, and the quality hit at that scale is not worth it for most workloads.
Ollama: there is no ollama pull qwen3.8-max that fits your machine. If you want a locally runnable Qwen, look at the smaller siblings in the family (Qwen3 30B-A3B and similar MoE variants) that target consumer VRAM budgets. Those are the models madebyagents.com covers in the local-hardware tier.
If your actual goal is local inference, Qwen3.8-Max is the wrong target. Use it through Alibaba Cloud Model Studio as a reference or teacher model, and run a smaller Qwen locally.
vs Kimi K3 (Moonshot): Kimi K3 is open-weight and launched days before Qwen3.8-Max's preview. If open weights ship on schedule, Qwen3.8-Max competes on context length (1M) and multimodal breadth; Kimi K3 competes on openness and licensing clarity. Both are datacenter models, not local ones.
vs Claude Fable 5: Fable 5 leads on SWE-bench Pro (80.0 vs 67.7) and HLE (53.3 vs 43.6). Qwen3.8-Max wins on PaperBench (93.0) and IFBench (82.8), and it is cheaper at $2/$6 per million tokens. Choose Qwen3.8-Max when cost and 1M context dominate; choose Fable 5 when hard agentic coding accuracy is the priority.
For local deployments, neither is the answer. The comparison that matters for madebyagents.com readers is between the smaller Qwen MoE variants you can actually run, and those are a different page.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Alibaba model we track.

Explore the Family
The full Qwen family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.