MiMo-V2.6-Flash-RL is an open-weights model from the Xiaomi MiMo team, released as the efficiency-balanced checkpoint of the MiMo-V2.6 series. It is a sparse mixture-of-experts model with 310.76B total and 15B active parameters, a 1M token context window, and text, image, video and audio input. Training used one mixed reinforcement learning run across coding, general agents, visual tasks and cybersecurity, with groupwise agentic grading for self-improvement. Weights are MIT licensed on Hugging Face, and the MiMo API platform lists $0.14 per million input tokens and $0.28 per million output tokens for the Flash model.
A solid 310.76B-parameter MoE language model from Xiaomi MiMo. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 126.0 GB | Low | |
| Q4_K_MRecommended | 129.2 GB | Good | |
| Q5_K_M | 130.7 GB | Very Good | |
| Q6_K | 132.5 GB | Excellent | |
| Q8_0 | 136.3 GB | Near Perfect | |
| FP16 | 150.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| Google TPU v7 (Ironwood)Google | SS | 46.0 tok/s | 129.2 GB |
| NVIDIA B200 GPUNVIDIA | SS | 49.8 tok/s | 129.2 GB |
| AMD Instinct MI355XAMD | SS | 49.8 tok/s | 129.2 GB |
| AMD Instinct MI325XAMD | SS | 37.4 tok/s | 129.2 GB |
| AMD Instinct MI300XAMD | SS | 33.0 tok/s | 129.2 GB |
Energy cost on Apple M4 Max (40-core GPU) (~3.4 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)MiMo-V2.6-Flash-RL on Apple M4 Max (40-core GPU) · ~3.4 tok/s · 92W | $0.901 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 129 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
AMD Instinct MI300XRunPod · Community · 192 GB VRAM | $0.50 |
AMD Instinct MI300XRunPod · Spot · 192 GB VRAM | $0.50 |
NVIDIA H200 NVLRunPod · Community · 141 GB VRAM | $0.50 |
NVIDIA H200 NVLRunPod · Spot · 141 GB VRAM | $0.50 |
NVIDIA H200 SXMVast.ai · Spot · 141 GB VRAM | $1.32 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Xiaomi MiMo model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.