K2-Horizon-7B is a 7B-parameter dense decoder-only model from the Institute of Foundation Models (IFM), the medium dense member of the six-model K2 Horizon fleet. It is built for agentic, coding, long-context and reasoning tasks, with a native 524,288-token context window and intermediate checkpoints released so capability growth can be studied. Weights, training code, data recipes and evaluation resources are published under the Apache 2.0 license, and vLLM and SGLang serving recipes ship with reasoning and tool-call parsers.
A workable 7B-parameter dense language model from Institute of Foundation Models (IFM). Pulls ahead on graduate-level reasoning (GPQA) (75/100), so reach for it when that's the dimension that matters.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 32.5 GB | Low | |
| Q4_K_MRecommended | 33.9 GB | Good | |
| Q5_K_M | 34.6 GB | Very Good | |
| Q6_K | 35.5 GB | Excellent | |
| Q8_0 | 37.2 GB | Near Perfect | |
| FP16 | 43.9 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| NVIDIA A100 SXM4 80GBNVIDIA | SS | 48.4 tok/s | 33.9 GB |
| NVIDIA H100 SXM5 80GBNVIDIA | SS | 79.5 tok/s | 33.9 GB |
| Google Cloud TPU v5pGoogle | SS | 65.6 tok/s | 33.9 GB |
| Intel Gaudi 2 AI AcceleratorIntel | SS | 58.1 tok/s | 33.9 GB |
| Intel Gaudi 3 AI AcceleratorIntel | AA | 87.8 tok/s | 33.9 GB |
Energy cost on Apple Mac Mini (M4, 2024) (~2.8 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)K2-Horizon-7B on Apple Mac Mini (M4, 2024) · ~2.8 tok/s · 55W | $0.644 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 34 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA RTX PRO 5000 BlackwellVast.ai · Spot · 48 GB VRAM | $0.27 |
NVIDIA A100 80GB PCIeVast.ai · Spot · 80 GB VRAM | $0.29 |
NVIDIA RTX A6000RunPod · Community · 48 GB VRAM | $0.33 |
NVIDIA RTX A6000RunPod · Spot · 48 GB VRAM | $0.33 |
NVIDIA A40RunPod · Community · 48 GB VRAM | $0.35 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
K2-Horizon-7B is a 7B-parameter dense decoder-only model developed by the Institute of Foundation Models (IFM). Positioned as the medium dense engine in IFM's six-model K2 Horizon family, it targets practitioners looking for high-performance agentic autonomy, software engineering capabilities, and deep mathematical reasoning on local hardware. Unlike many open-weight releases that keep post-training pipelines proprietary, IFM provides full transparency under an Apache 2.0 license, publishing the complete training pipeline, data recipes, intermediate checkpoints, and serving configurations.
Despite sitting in the standard 7B weight class, K2-Horizon-7B redefines what developers can expect from a local AI model with 7B parameters in 2026. Out of the box, it features an architectural context window of 524,288 tokens (512K) established during midtraining. Benchmark evaluations place it ahead of several larger sub-15B competitors, showing exceptional proficiency on complex coding pipelines, long-horizon tool execution, and competition-level mathematics.
For engineers evaluating self-hosted models, K2-Horizon-7B fills the gap between lightweight utility models and heavy MoE networks. It provides the low inference latency and straightforward deployment characteristics of a dense 7B architecture while delivering reasoning scores that directly challenge models nearly twice its size.
K2-Horizon-7B uses a refined dense autoregressive transformer architecture. By opting for a standard dense design rather than a Mixture of Experts (MoE), IFM ensures predictable memory bandwidth utilization and consistent compute throughput across all token generation stages.
The defining technical attribute of the model is its native 524,288-token context window. Rather than relying solely on post-hoc RoPE scaling at inference time, IFM integrated this extended context across its midtraining stages. This minimizes attention degradation and needle-in-a-haystack retrieval loss across deep prompt depths.
However, running full 512K context prompts locally requires careful consideration of the Key-Value (KV) cache:
IFM also validated diffusion-based adapters (such as K2-Horizon-7B-Uno) for speculative decoding pipelines, allowing developers to compress decode latency during memory-bound local execution without destabilizing output quality.
K2-Horizon-7B is tuned specifically for agentic environments, technical problem-solving, and code synthesis. Rather than focusing purely on open-ended conversational chat, the model emphasizes structured deterministic execution and rigorous chain-of-thought processing.
The model sets a high bar on software engineering benchmarks, scoring 70.6% on SWE-bench Verified and 39.1% on Terminal-Bench 2.1. This makes it an exceptional engine for autonomous developer agents. Concrete applications include:
With scores of 25.8% on tau3-Banking and 59.0% on BrowseComp, K2-Horizon-7B behaves reliably inside agentic harnesses. It respects strict JSON schemas, outputs well-formed API payloads, and maintains task state across dozens of sequential tool interactions. It excels in local enterprise workflows requiring autonomous web search synthesis, database query construction, and automated ticket triage.
On competition-level mathematics, K2-Horizon-7B recorded a 73.3% score on HMMT Feb 2026, alongside an 18.6% on the Humanity's Last Exam (HLE) scientific benchmark and 31.6% on SciCode. These results make it an ideal local copilot for scientific computing, algorithmic problem solving, and formal logic verification where cloud data privacy cannot be compromised.
To run K2-Horizon-7B locally, practitioners need to plan around two variables: model weight storage and KV cache allocation for their target context size.
The weights themselves fit easily on standard consumer hardware, but context expansion determines your overall VRAM ceiling:
Note on 512K Context: Utilizing the entire 524,288 context window requires substantial VRAM for the KV cache. To run the full 512K context locally, you will need a dual-GPU setup (such as 2x RTX 3090/4090 24GB with FP8 KV cache) or a unified memory system such as a Mac Studio with 64 GB to 128 GB of RAM.
For everyday local deployment, Q4_K_M is the best quantization for K2-Horizon-7B. It preserves over 98% of the baseline reasoning capability on SWE-bench and math evaluations while dropping weight overhead down to roughly 5.2 GB. If your workload involves dense scientific logic or competitive math, step up to Q8_0 or FP8 if VRAM permits.
K2-Horizon-7B tokens per second vary by runtime framework and hardware:
The quickest way to get started with the model locally is via Ollama:
1ollama run ifm/k2-horizon-7b
For agentic tool use and maximum throughput, deploy the model through vLLM:
1vllm serve IFM/K2-Horizon-7B \2 --max-model-len 32768 \3 --gpu-memory-utilization 0.95 \4 --enable-reasoning \5 --tool-call-parser hermes
Evaluating K2-Horizon-7B performance against common alternatives in the 8B to 12B range demonstrates how aggressively IFM tuned this architecture:
Qwen3.5-9B has been a standard benchmark in this tier, but K2-Horizon-7B outperforms it decisively across critical engineering vectors:
Google's Gemma 4-12B requires roughly 40% more VRAM to run at equivalent precision levels, yet falls behind K2-Horizon-7B on key technical tasks:
For developers seeking an unencumbered, Apache 2.0 licensed model that runs comfortably on a single consumer GPU while outperforming 9B-12B alternatives on code and math, K2-Horizon-7B is one of the strongest 7B deployments currently available.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Institute of Foundation Models (IFM) model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.