K2-Horizon-MoVA-36B-A4B is a sparse mixture-of-experts model from the Institute of Foundation Models, released under Apache 2.0. It stores 36B parameters and runs 4B per token, using a Mixture-of-Values attention (MoVA) design. Native context is 524,288 tokens from the midtraining stages onward. It is aimed at agentic tool use, coding and reasoning, and the final checkpoint is published on Hugging Face with training data and code to follow.
A workable 36B-parameter MoE language model from Institute of Foundation Models. Pulls ahead on graduate-level reasoning (GPQA) (82/100), so reach for it when that's the dimension that matters. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 18.8 GB | Low | |
| Q4_K_MRecommended | 19.6 GB | Good | |
| Q5_K_M | 20.0 GB | Very Good | |
| Q6_K | 20.5 GB | Excellent | |
| Q8_0 | 21.5 GB | Near Perfect | |
| FP16 | 25.3 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| Google Cloud TPU v6e (Trillium)Google | SS | 67.4 tok/s | 19.6 GB |
| NVIDIA GeForce RTX 5090 Founders EditionNVIDIA | SS | 73.6 tok/s | 19.6 GB |
| Origin PC M-CLASS v2Origin PC | SS | 73.6 tok/s | 19.6 GB |
| NVIDIA RTX 6000 Ada GenerationNVIDIA | SS | 39.4 tok/s | 19.6 GB |
| Origin PC L-CLASS v2Origin PC | SS | 39.4 tok/s | 19.6 GB |
Energy cost on AMD Radeon RX 7900 XT (~33 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)K2-Horizon-MoVA-36B-A4B on AMD Radeon RX 7900 XT · ~33 tok/s · 315W | $0.320 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 20 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.11 |
NVIDIA RTX PRO 4000 BlackwellVast.ai · Spot · 24 GB VRAM | $0.13 |
NVIDIA A100 80GB SXMVast.ai · Spot · 80 GB VRAM | $0.13 |
NVIDIA GeForce RTX 3090Vast.ai · On-Demand · 24 GB VRAM | $0.14 |
NVIDIA RTX A5000RunPod · Community · 24 GB VRAM | $0.16 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
K2-Horizon-MoVA-36B-A4B is a sparse mixture-of-experts model from the Institute of Foundation Models, released under Apache 2.0. It stores 36B parameters and activates roughly 4B per token, which puts it in a category that barely existed two years ago: models that fit on a single high-end workstation while producing output quality that open dense models needed three to four times the parameter count to reach. The attention design is not standard multi-head attention. IFM uses Mixture-of-Values attention (MoVA), which is the architectural detail that separates this checkpoint from the rest of the 30B-class field.
The pitch is narrow and specific. This is not a general-purpose chat model competing on vibes. IFM trained it for agentic tool use, coding, and reasoning, and the published numbers back that up. On tau3-Banking agentic tool use it scores 26.8, ahead of Nemotron 3 Ultra (550B total, 55B active) at 14.2 and Muse Glimmer-30B at 23.5. On Terminal-Bench 2.1 it posts 58.6 against 53.9 for Nemotron 3 Ultra and 44.9 for Qwen3.6-35B-A3B. GPQA Diamond lands at 80.8, within a few points of dense models twice its active footprint.
The 524,288-token native context is the other headline. That is 512K, available from the midtraining stages onward rather than bolted on with a rope-scaling patch after the fact. For anyone building document ingestion, repo-scale code analysis, or long-horizon agent loops, that changes what is feasible on local hardware.
The MoE design is the reason this model is interesting for local deployment, and also the reason its VRAM math confuses people.
Active parameters govern compute, not memory. Every one of the 36B weights must be resident in VRAM or system RAM because the router can select any expert on any token. What you save is arithmetic: the model does roughly the work of a 4B dense network per forward pass. In practice that means decode throughput closer to a 4B model than a 36B one, while the memory footprint stays at 36B.
MoVA replaces the standard attention value projection with a mixture over value subspaces. IFM has not published the full routing detail in the model card, but the practical effect claimed is better quality per active parameter at long context, which is where attention variants usually fall apart.
That 512K window is not free. KV cache scales linearly with context, and at 524K tokens a full-precision cache will exhaust any consumer card long before the weights do. Budget for quantized KV cache (q8_0 or q4_0 in llama.cpp terms) or cap the working context at 32K-128K for most workloads and reserve the full window for batch jobs.
The declared capabilities are chat, code, reasoning, function-calling, and math. Where it earns its keep:
It is text-only. No vision, no audio. If your pipeline needs image input, this is not the model.
The weights are on Hugging Face under IFM/K2-Horizon-MoVA-36B-A4B, with intermediate checkpoints tagged (base_final, mid_1_final, mid_4_final) for staged comparison. Training data and code are slated for release.
VRAM requirements by quantization:
Realistic hardware:
Getting started: the fastest path is Ollama or llama.cpp with a GGUF build. Pull a Q4_K_M quant, set num_ctx to your actual working window rather than the full 524K, and enable KV cache quantization. vLLM and SGLang both have serving recipes in the model's technical appendix if you need throughput or continuous batching.
Quantization guidance: start at Q4_K_M. Move to Q5_K_M or Q8_0 only if your evals show measurable degradation on reasoning tasks. MoE models are generally more robust to aggressive quantization than dense models because routing decisions are less sensitive to small weight perturbations, but the 4B active budget leaves less margin than a 70B model would have.
vs. Qwen3.6-35B-A3B. Nearly identical parameter class (35B total, 3B active) and the closest direct competitor. K2-Horizon leads on agentic benchmarks by a wide margin (tau3-Banking 26.8 vs 9.3, Terminal-Bench 58.6 vs 44.9) and on GPQA Diamond. Qwen3.6 has a larger ecosystem of fine-tunes and better tooling maturity. Choose Qwen if you need community support and broad fine-tune availability; choose K2-Horizon if agentic reliability is the priority.
vs. Gemma 4 31B-it (dense). Gemma is dense, so all 31B parameters compute on every token. It is slower to decode and heavier on memory per unit of quality, but simpler to serve and easier to reason about. K2-Horizon wins on tool use and terminal tasks; Gemma holds its own on scientific reasoning (85.7 on GPQA Diamond vs 80.8). If your workload is single-turn QA rather than agents, the dense model is a defensible pick.
The Apache 2.0 license is the tiebreaker for commercial deployment. No gating, no usage restrictions, no per-seat fees.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Institute of Foundation Models model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.