Kimi K3 is Moonshot AI's flagship open-weight model, announced in July 2026. It is a 2.8 trillion parameter mixture-of-experts model with 104B active parameters, native vision and a 1 million token context window. It is built for long-horizon coding, agentic knowledge work and reasoning, and runs on Kimi.com, Kimi Work, Kimi Code and the Kimi API. Full weights are released under the Kimi K3 License.
A solid 2800B-parameter MoE language model from Moonshot AI. Pulls ahead on graduate-level reasoning (GPQA) (94/100), so reach for it when that's the dimension that matters.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 911.4 GB | Low | |
| Q4_K_MRecommended | 933.2 GB | Good | |
| Q5_K_M | 943.6 GB | Very Good | |
| Q6_K | 956.1 GB | Excellent | |
| Q8_0 | 982.1 GB | Near Perfect | |
| FP16 | 1080.9 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | FF | 0.4 tok/s | 933.2 GB |
| Acer Veriton GN100 AI MiniAcer | FF | 0.2 tok/s | 933.2 GB |
| AMD Instinct MI300XAMD | FF | 4.6 tok/s | 933.2 GB |
| AMD Instinct MI325XAMD | FF | 5.2 tok/s | 933.2 GB |
| AMD Instinct MI355XAMD | FF | 6.9 tok/s | 933.2 GB |
Energy cost running this model locally vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
| No featured GPU in our directory fits this model. Check back as we add more hardware. | |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 933 GB VRAM, refreshed hourly.
No current rental listing covers this model’s VRAM requirement on the providers we track.
Kimi K3 is Moonshot AI's flagship open-weight model, announced in July 2026. It is a 2.8 trillion parameter mixture-of-experts model with 104B active parameters, native vision, and a 1,048,576 token context window. Moonshot calls it the world's first open 3T-class model. It sits just below Claude Fable 5 and GPT 5.6 Sol on most benchmarks, and ahead of Claude Opus 4.8 and GPT 5.5. For practitioners evaluating local AI hardware, Kimi K3 is a datacenter-class model: it demands serious infrastructure, not a consumer GPU.
Kimi K3 is built for long-horizon coding, agentic knowledge work, and reasoning. It runs on Kimi.com, Kimi Work, Kimi Code, and the Kimi API. The full weights are released under the Kimi K3 License. Unlike smaller open models that target single-GPU workflows, Kimi K3 targets multi-GPU servers and high-memory workstations. If you are searching for how to run a 2800B parameter local AI model in 2026, this is the reality check.
This model competes with proprietary frontier systems but gives you open weights. That matters for research, fine-tuning, and on-prem deployment where API access is not an option. It also means you own the inference stack. The tradeoff is hardware: Kimi K3 VRAM requirements are measured in terabytes, not gigabytes. We will cover what that means for quantization, tokens per second, and whether any consumer hardware can touch it.
Kimi K3 uses a Mixture-of-Experts (MoE) architecture with 2800B total parameters and 104B active parameters. That distinction is the most important thing to understand for local deployment. Compute cost per token is roughly equivalent to a 104B dense model. Memory footprint is that of a 2800B model. You must hold all expert weights in VRAM or system RAM, even though only a fraction activate per token.
The MoE sparsity is aggressive: 16 out of 896 experts activate per token. Moonshot paired this with a Stable LatentMoE framework. The model also introduces Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two architectural changes that improve information flow across sequence length and model depth. Moonshot claims approximately 2.5x improvement in scaling efficiency over Kimi K2.
The context window is 1,048,576 tokens. That enables whole-repository code analysis, hour-long video understanding, and multi-document research without chunking. The caveat is KV cache. At 1M tokens, the KV cache alone can consume tens of gigabytes, which compounds your VRAM requirements beyond the model weights.
Native multimodality means text, images, and video are processed by the same model. There is no separate vision encoder to manage. For inference speed, the active parameter count (104B) sets the compute ceiling. The total parameter count (2800B) sets the memory bandwidth floor. Kimi K3 MoE efficiency is excellent when all experts reside in fast memory, but it becomes memory-bandwidth bound the moment you offload experts to CPU RAM.
Kimi K3 excels at long-horizon coding. It sustains engineering sessions with minimal human oversight, navigates massive repositories, and orchestrates terminal tools. Concrete examples from Moonshot's launch materials include GPU kernel optimization, compiler development, vision-in-the-loop game development, CAD work, and chip design. In one demonstration, Kimi K3 designed and verified a prototype chip in 48 hours with no human intervention, closing timing at 100 MHz and hitting over 8,700 tokens per second in simulation.
Agentic knowledge work is the second major strength. The model produces deep research with interactive visualizations, widgets, and dashboards. It handles motion design and video editing. Function calling is supported, which makes it suitable for autonomous agents that need to invoke APIs, query databases, or control software. Vision capabilities let it read screenshots, diagrams, charts, and video frames, so it can operate GUI-based tools or verify visual output.
Reasoning benchmarks back the claims. On WebDev Arena, Kimi K3 became the first open model to top the leaderboard with 1,678 Elo, ahead of Claude Fable 5 at 1,634. On BrowseComp, it scored 91.2% at $2.03 per task, roughly half the cost of GPT-5.6 Sol and an order of magnitude less than Claude models. For Kimi K3 for coding, the model is competitive with the best proprietary systems. For Kimi K3 reasoning benchmark performance, it consistently outperforms other open models.
Practical use cases include:
Local deployment of Kimi K3 is not for typical consumer hardware. The Kimi K3 hardware requirements start in the hundreds of gigabytes for even the most aggressive quantization. Here are rough VRAM estimates based on parameter count:
These are approximations. Actual quantized sizes vary with the quantization method and any compressed-tensor formats. The best quantization for Kimi K3 depends on your memory budget. Q4_K_M is the recommended balance for most users who have ~1.5TB of VRAM or unified memory. Q2_K is available for extreme setups, but quality degradation is noticeable on reasoning and coding tasks.
No consumer GPU can run Kimi K3 at usable speeds. An RTX 4090 has 24GB of VRAM. An M4 Max tops out at 128GB of unified memory. Even eight RTX 5090 cards (32GB each, 256GB total) are not enough for 4-bit weights. You need datacenter accelerators. For 4-bit, eight MI300X GPUs (192GB each, 1.536TB total) are a realistic minimum. For 8-bit, sixteen H200 GPUs (141GB each, 2.256TB total) are required. The best GPU for Kimi K3 in 2026 is the H200 or B200 in multi-GPU configurations.
How to run a 2800B model on a consumer GPU? You cannot, unless you accept CPU offloading with 1TB or more of system RAM plus a single GPU for compute. That setup yields roughly 1-5 tokens per second. Kimi K3 tokens per second on full GPU residency is better, typically 20-50 tokens per second depending on batch size and context length. The 104B active parameters keep compute manageable, but memory bandwidth is the bottleneck.
Ollama is not a practical way to run Kimi K3. Ollama works well for 7B-70B models, but a 2.8T MoE model requires a multi-GPU server. Use vLLM or llama.cpp with tensor parallelism on a dedicated inference node. If you only have consumer hardware, look at smaller MoE models instead. Kimi K3 is an open-weight model, but open weights do not mean accessible weights.
There is no direct open-weight competitor at 2800B parameters. The closest alternatives are smaller. DeepSeek-V3 is a 671B parameter MoE with 37B active parameters. It is roughly four times smaller in total parameters. At 4-bit, DeepSeek-V3 needs around 400GB of VRAM, which fits on a single 8x A100 80GB node. Llama 3.1 405B is a dense model. At 4-bit, it needs around 200GB of VRAM, which fits on 4x A100 80GB. Both are far more practical for local deployment than Kimi K3.
Kimi K3 outperforms both on most reasoning, coding, and agentic benchmarks. It also has a much larger context window (1,048,576 tokens vs 128K for DeepSeek-V3 and Llama 3.1). The tradeoff is hardware scale. If you have a single 8-GPU node, DeepSeek-V3 gives you strong performance with lower memory pressure. If you have a 16-GPU or larger cluster, Kimi K3 delivers frontier-level open weights.
Choose Kimi K3 when you need the best open-weight model for long-horizon coding, multimodal reasoning, and agentic knowledge work, and you have datacenter infrastructure to match. Choose DeepSeek-V3 when you want a capable MoE that runs on a standard 8-GPU node. Choose Llama 3.1 405B when you prefer a dense architecture with broad ecosystem support. For anyone with a single RTX 4090 or M4 Max, none of these models run locally at acceptable speeds. Look at 70B or smaller models for consumer hardware.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Moonshot AI model we track.

Explore the Family
The full Kimi family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.