JetBrains released Mellum2.1 on 2026-10-08, a 12B mixture-of-experts coding model with 2.5B active parameters. It is built for coding agents and fast sub-agents that run on your own hardware, and it can explore a codebase, edit files and check its own changes. Training was mostly reinforcement learning in thousands of sandboxed environments, with added tasks in math, competitive programming, science, tool use and software engineering. Weights are open under the Apache 2.0 license.
A solid 12B-parameter MoE language model from JetBrains. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 1.5 GB | Low | |
| Q4_K_MRecommended | 2.0 GB | Good | |
| Q5_K_M | 2.3 GB | Very Good | |
| Q6_K | 2.6 GB | Excellent | |
| Q8_0 | 3.2 GB | Near Perfect | |
| FP16 | 5.6 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7600 8GBAMD | SS | 114.1 tok/s | 2.0 GB |
| NVIDIA GeForce RTX 4060NVIDIA | SS | 107.8 tok/s | 2.0 GB |
| NVIDIA GeForce RTX 5060 Ti 8GBNVIDIA | SS | 177.5 tok/s | 2.0 GB |
| NVIDIA Jetson Orin Nano Super Developer KitNVIDIA | SS | 40.4 tok/s | 2.0 GB |
| AMD Radeon RX 7700 XTAMD | SS | 171.1 tok/s | 2.0 GB |
Energy cost on Raspberry Pi 5 (8GB) (~13 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Mellum2.1 on Raspberry Pi 5 (8GB) · ~13 tok/s · 12W | $0.030 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 2 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.07 |
NVIDIA GeForce RTX 4060Vast.ai · On-Demand · 8 GB VRAM | $0.08 |
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.08 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
JetBrains released Mellum2.1 on October 8, 2026, delivering an open-weights coding model engineered specifically for autonomous software agents and local execution. Built as a 12B parameter Mixture-of-Experts (MoE) network with only 2.5B active parameters per token, Mellum2.1 delivers the execution throughput of a sub-3B model while maintaining the reasoning capacity of a much larger dense network. The weights are distributed under an Apache 2.0 license, making it fully viable for commercial workflows, private on-premises code review, and embedded development tools.
Unlike generalist models that struggle with multi-step terminal interactions, Mellum2.1 targets the agentic loop: exploring a repository, planning changes, editing source files, and verifying modifications through shell commands and test runners. JetBrains developed the model primarily through extensive reinforcement learning inside millions of sandboxed environments. This focus resolves the core failure mode of earlier lightweight coding models, which often generated syntactically valid code snippets in isolation but failed when navigating real-world directory trees and build systems.
For engineers seeking a local AI model with 12B parameters in 2026, Mellum2.1 represents a practical shift toward specialized agent workers. It avoids the latency penalties and cloud privacy liabilities associated with hosted APIs, offering developers a dedicated, private engine for fast sub-agent execution, continuous test repair, and automated code review directly on consumer-grade hardware.
The core strength of Mellum2.1 lies in its sparse Mixture-of-Experts design. While the model contains 12B parameters in total memory, the routing layer activates only 2.5B parameters for any given forward pass. This split delivers significant advantages for local deployment:
Understanding Mellum2.1 MoE efficiency requires separating memory capacity from compute throughput. Memory bandwidth and capacity dictate that your system must hold the full 12B parameters, requiring between 8 GB and 24 GB of memory depending on your quantization level. However, during token generation, the compute engine only evaluates 2.5B parameters per token. This low computational footprint translates to high generation speeds, even on memory-constrained GPUs or unified-memory Apple Silicon systems where raw compute is typically the bottleneck.
The native 131K context window is particularly relevant for software engineering tasks. It allows an agent to ingest full file trees, entire API documentation sets, and substantial debugging traces without forcing aggressive context truncation or relying solely on brittle retrieval heuristics.
Mellum2.1 was trained using reinforcement learning against verifiable outcomes: passing unit tests, valid bash execution, and correct math proofs. Consequently, it excels across several distinct local engineering workflows:
Mellum2.1 functions effectively inside agent frameworks such as Aider, JetBrains Air, or custom CLI loops. It evaluates directory structures, determines which files require edits, applies precise unified diffs, and inspects terminal error logs to fix syntax or runtime errors iteratively. On the SWE-bench Verified benchmark, Mellum2.1 reaches a 47.0% resolve rate, jumping significantly over previous iterations.
With a Berkeley Function-Calling Leaderboard (BFCL v4) score of 62.3%, the model handles structured JSON tool calls, database queries, and custom shell operations reliably. It knows when to emit raw shell commands, when to halt for tool output, and how to parse returned execution frames.
On core reasoning benchmarks, the model scores 82.0% on LiveCodeBench v6, 91.5% on HumanEval+, and 83.3% exact match on AIME 25/26. The integrated thinking mode allows Mellum2.1 to generate internal reasoning traces before returning code, making it particularly capable at complex algorithmic problems, dynamic programming challenges, and architectural refactoring.
Running Mellum2.1 locally gives development teams zero-latency code generation without data leakage. Because it is an MoE architecture, hardware planning requires enough VRAM to hold the 12B total parameters while sizing compute expectations around a fast 2.5B core.
Memory requirements depend on the chosen precision and the length of your active context window. At deep context lengths (up to 131K), the Key-Value (KV) cache consumes additional memory:
For most local practitioners, the best quantization for Mellum2.1 is Q4_K_M. It preserves the vast majority of the model's agentic and coding performance while keeping the base weight footprint under 8 GB.
Because of its 2.5B active parameter compute load, Mellum2.1 tokens per second figures significantly outperform standard dense 12B or 14B models:
The quickest way to get started with the model locally is via Ollama or vLLM:
1# Using vLLM with the thinking parser2vllm serve JetBrains/Mellum2.1-12B-A2.5B-Thinking --reasoning-parser qwen3 --max-model-len 3276834# Using Ollama (once published or imported from GGUF)5ollama run mellum2.1:12b
When serving Mellum2.1 in high-throughput local pipelines, using vLLM allows you to leverage continuous batching and multi-token prediction heads where supported.
To evaluate Mellum2.1 for coding pipelines, consider how it stacks up against similar small-footprint models:
Qwen3.5-9B is a dense model that delivers high baseline general reasoning and slightly edges out Mellum2.1 on raw SWE-bench Verified scores (roughly 50.0% versus Mellum2.1's 47.0%). However, Mellum2.1 holds two practical advantages for local deployments:
DeepSeek-Coder-V2-Lite uses an MoE design with 2.4B active parameters out of 16B total. While its base architecture is broadly comparable in active compute, Mellum2.1 benefits from a more compact overall memory footprint (12B total parameters vs 16B), requiring less VRAM to store the weight tables. Furthermore, JetBrains trained Mellum2.1 specifically inside modern tool-use sandboxes, resulting in stronger out-of-the-box reliability for unified terminal execution, shell tool use, and iterative patch verification.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every JetBrains model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.