Kolibri is an open-weight Mixture-of-Experts language model from Aleph Alpha in Heidelberg, released on 3 October 2026 under the Apache 2.0 license. It has 78B total parameters with 3B active and supports context lengths up to 1M tokens. Aleph Alpha built it for sovereign, mission-critical work in regulated sectors such as public administration, industrials and aerospace, and specialized it for German, reasoning, math and agentic behavior. Full weights can be downloaded from Hugging Face and served with the company's vLLM plugin.
A solid 78B-parameter MoE language model from Aleph Alpha. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 25.6 GB | Low | |
| Q4_K_MRecommended | 26.2 GB | Good | |
| Q5_K_M | 26.5 GB | Very Good | |
| Q6_K | 26.9 GB | Excellent | |
| Q8_0 | 27.6 GB | Near Perfect | |
| FP16 | 30.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| NVIDIA A100 SXM4 80GBNVIDIA | SS | 62.6 tok/s | 26.2 GB |
| NVIDIA H100 SXM5 80GBNVIDIA | SS | 102.8 tok/s | 26.2 GB |
| Google Cloud TPU v5pGoogle | SS | 84.8 tok/s | 26.2 GB |
| Intel Gaudi 2 AI AcceleratorIntel | SS | 75.2 tok/s | 26.2 GB |
| Intel Gaudi 3 AI AcceleratorIntel | SS | 113.5 tok/s | 26.2 GB |
Energy cost on AMD Radeon RX 7900 XTX (~29 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Kolibri on AMD Radeon RX 7900 XTX · ~29 tok/s · 355W | $0.402 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 26 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA RTX A6000RunPod · Community · 48 GB VRAM | $0.33 |
NVIDIA RTX A6000RunPod · Spot · 48 GB VRAM | $0.33 |
NVIDIA A40RunPod · Community · 48 GB VRAM | $0.35 |
NVIDIA A40RunPod · Spot · 48 GB VRAM | $0.35 |
NVIDIA A40RunPod · Secure · 48 GB VRAM | $0.49 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
Kolibri is an open-weight, text-only language model from Aleph Alpha, the Heidelberg lab known for sovereign AI deployments in European public administration and industry. Aleph Alpha released it on 3 October 2026 under the Apache 2.0 license, with full weights on Hugging Face and an official vLLM plugin for serving. It is a Mixture-of-Experts transformer with 78B total parameters and roughly 3B active per token, which puts it in a narrow category: models with large total capacity that stay cheap to compute per token.
The model's pitch is specific. It is built for German and English, tuned for reasoning, math, and agentic tool calling, and supports a context window that Aleph Alpha validates up to 1,048,576 tokens. That combination targets regulated, mission-critical workloads where data cannot leave a controlled environment and where documents are long, structured, and frequently German. If you need a general-purpose English assistant, plenty of open models do that well. If you need a German-first model with a million-token context and permissive licensing, the field is much thinner.
Kolibri competes most directly with Llama 3.3 70B and the open-weight MoE releases in the 100B class, such as OpenAI's gpt-oss-120B and Qwen3's larger variants. Its differentiators are the German specialization, the validated long context, and Apache 2.0 terms that carry no usage restrictions or regional carve-outs.
Kolibri is a 50-layer MoE transformer with 384 experts per layer, one shared expert and six routed per token, using a 4:1 sliding-window to grouped-query attention ratio. Training used Muon and Exact Quantile Balancing. Aleph Alpha trained it on 20T tokens (roughly 62.5% English, 23.9% German, 13.6% code), followed by 3.44T mid-training tokens and a 201B-token long-context extension phase. The full pre-training run took 21 days on 768 NVIDIA B200 GPUs.
What active parameters mean in practice: only about 3.46B of the 78B parameters are used for each token generated. Inference compute therefore resembles a small dense model, while the memory footprint remains that of a 78B model, because every expert must be resident or quickly paged in. This is the central Kolibri MoE efficiency tradeoff. You get fast decode relative to parameter count, but you still need the VRAM to hold the full expert set.
Weights ship in float8_e4m3fn at 128×128 block granularity with dynamically quantized activations and an FP8 KV cache. Embeddings, LM head, norms, and the MoE router stay in bfloat16. That puts the FP8 memory footprint at roughly 78 GB.
Context is worth understanding precisely. The native trained context is 262,144 tokens. Positional encoding is applied only in the sliding-window layers, so the context can extend beyond that without position scaling, in principle to arbitrary length. Aleph Alpha has validated quality and serving efficiency up to 1,048,576 tokens. For latency-sensitive or complex tasks, they recommend staying at or below 262,144 tokens. The model card lists a knowledge cutoff of 18 June 2026 for both languages, though tool use can pull in newer information.
Kolibri is tuned for German, reasoning, math, and agentic behavior, and it supports an explicit reasoning mode plus tool calling. Its strengths map onto concrete work:
Language coverage is the main limitation. Kolibri handles German and English. It is not a broad multilingual model, and you should not expect strong performance in French, Spanish, or Chinese. For those, Qwen3 or Llama variants are better picks.
The official serving path is vLLM with Aleph Alpha's plugin. That is what you should use for production, because it handles the FP8 format and the long-context KV cache correctly.
Memory requirements:
Q4_K_M: roughly 40-45 GB. The pragmatic choice for most local users.Q5_K_M: roughly 50-55 GB, a modest quality bump over Q4.Q8_0: roughly 80+ GB, close to FP8 quality but with no real memory savings.What consumer hardware can actually run it: an RTX 4090 with 24 GB cannot hold Kolibri alone. Two 4090s (48 GB) can run Q4_K_M with partial CPU offload, but throughput drops sharply. Apple silicon is the more realistic consumer path: an M4 Max with 128 GB or an M3 Ultra Mac Studio with 256 GB or more can hold Q4_K_M or Q5_K_M entirely in unified memory, though at lower token rates than datacenter GPUs. A 64 GB M4 Max will run Q4_K_M with a reduced context window.
Expected throughput (rough estimates, batch size 1, moderate context):
Q4_K_M: roughly 15-25 tokens/secQ4_K_M: roughly 8-18 tokens/secLong context changes these numbers. The KV cache scales with sequence length, and at 1M tokens even an FP8 KV cache is substantial. If you need the full window, budget memory for it explicitly.
Getting started: Ollama or llama.cpp with a GGUF build is the quickest route once a conversion is available, since it abstracts quantization and offload. For anything beyond a test, move to vLLM with the Aleph Alpha plugin. Check the Hugging Face repo for current GGUF availability before planning a deployment around it.
Kolibri vs Llama 3.3 70B. Both are roughly 70B-class open models. Llama 3.3 is dense, so all 70B parameters activate per token, making it heavier at inference than Kolibri's 3B active set. Llama is stronger in general English tasks and has a far larger ecosystem of fine-tunes and tooling. Kolibri wins on German, on context length (1M validated vs 128k), and on license terms (Apache 2.0 vs the Llama Community License). Choose Kolibri when German quality or long context is the deciding factor.
Kolibri vs gpt-oss-120B. OpenAI's gpt-oss-120B is also a MoE model with Apache 2.0 licensing, 117B total parameters, and roughly 5B active. It has a 131k context and is English-centric. It is a strong general reasoning and agentic model with wide runtime support. Kolibri is the better pick for German workloads and for contexts beyond 200k tokens; gpt-oss is the better pick for English-first agentic work where community tooling maturity matters more than context length.
Both comparisons point to the same decision rule. Kolibri is a specialized tool for German and English, long-context, sovereign deployments. If your workload matches that profile, the 78B memory footprint is the price of admission and the 3B active compute is what keeps it affordable to serve.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Aleph Alpha model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.