DeepSeek-V4.1-Flash is a multimodal Mixture-of-Experts model from DeepSeek-AI, released September 10, 2026. It has 552B backbone parameters and a Causal Encoder-Decoder architecture that activates 8B parameters per token during prefill and 16B during decode, with context up to one million tokens. Weights are open under the MIT license and the model is served on the DeepSeek API as deepseek-flash. It targets long-horizon agentic and input-heavy workloads, with a global KV cache of 890 bytes per token.
A solid 552B-parameter MoE language model from deepseek-ai. Pulls ahead on graduate-level reasoning (GPQA) (91/100), so reach for it when that's the dimension that matters. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 134.4 GB | Low | |
| Q4_K_MRecommended | 137.8 GB | Good | |
| Q5_K_M | 139.4 GB | Very Good | |
| Q6_K | 141.3 GB | Excellent | |
| Q8_0 | 145.3 GB | Near Perfect | |
| FP16 | 160.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| Google TPU v7 (Ironwood)Google | SS | 43.1 tok/s | 137.8 GB |
| NVIDIA B200 GPUNVIDIA | SS | 46.7 tok/s | 137.8 GB |
| AMD Instinct MI355XAMD | SS | 46.7 tok/s | 137.8 GB |
| AMD Instinct MI325XAMD | SS | 35.1 tok/s | 137.8 GB |
| ASUS ExpertCenter Pro ET900N G3ASUS | SS | 41.5 tok/s | 137.8 GB |
Energy cost on Apple M4 Max (40-core GPU) (~3.2 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)DeepSeek-V4.1-Flash on Apple M4 Max (40-core GPU) · ~3.2 tok/s · 92W | $0.961 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 138 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
AMD Instinct MI300XRunPod · Community · 192 GB VRAM | $0.50 |
AMD Instinct MI300XRunPod · Spot · 192 GB VRAM | $0.50 |
NVIDIA H200 NVLRunPod · Community · 141 GB VRAM | $0.50 |
NVIDIA H200 NVLRunPod · Spot · 141 GB VRAM | $0.50 |
NVIDIA H200 SXMVast.ai · Spot · 141 GB VRAM | $1.32 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every deepseek-ai model we track.

Explore the Family
The full DeepSeek family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.