Apertus 8B is an open-weights language model from EPFL, ETH Zurich and the Swiss National Supercomputing Centre (CSCS). It is the smaller of two Apertus releases, alongside a 70B version, and was trained on 15 trillion tokens across more than 1,000 languages, with 40 percent non-English data. Weights, training data and recipes are openly documented, and it is released under Apache 2.0. It is available via Hugging Face and the Swiss partner Swisscom.
A workable 8B-parameter dense language model from EPFL, ETH Zurich, CSCS. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 3.7 GB | Low | |
| Q4_K_MRecommended | 5.4 GB | Good | |
| Q5_K_M | 6.2 GB | Very Good | |
| Q6_K | 7.2 GB | Excellent | |
| Q8_0 | 9.2 GB | Near Perfect | |
| FP16 | 16.8 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7600 8GBAMD | SS | 42.9 tok/s | 5.4 GB |
| NVIDIA GeForce RTX 4060NVIDIA | SS | 40.5 tok/s | 5.4 GB |
| NVIDIA GeForce RTX 5060 Ti 8GBNVIDIA | SS | 66.8 tok/s | 5.4 GB |
| AMD Radeon RX 7700 XTAMD | SS | 64.4 tok/s | 5.4 GB |
| Intel Arc B580Intel | SS | 68.0 tok/s | 5.4 GB |
Energy cost on Raspberry Pi 5 (8GB) (~5.1 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Apertus 8B on Raspberry Pi 5 (8GB) · ~5.1 tok/s · 12W | $0.079 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 5 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3070RunPod · Community · 8 GB VRAM | $0.13 |
NVIDIA GeForce RTX 3070RunPod · Spot · 8 GB VRAM | $0.13 |
NVIDIA RTX A5000RunPod · Community · 24 GB VRAM | $0.16 |
NVIDIA RTX A5000RunPod · Spot · 24 GB VRAM | $0.16 |
NVIDIA GeForce RTX 3080RunPod · Community · 10 GB VRAM | $0.17 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Apertus 8B is an 8-billion-parameter dense language model from EPFL, ETH Zurich, and the Swiss National Supercomputing Centre (CSCS), released in September 2025 under Apache 2.0. It is the smaller half of a two-model family (the other is a 70B sibling) and the one that fits on hardware you probably already own. What separates it from most open-weight releases is the scope of the openness: weights, training data, filtering pipeline, training code, and intermediate checkpoints are all published. You can inspect and reproduce the run, not just download the artifact.
The differentiator is multilingual coverage at a scale no comparable 8B model matches. Pretraining ran over 15 trillion tokens across more than 1,000 languages, with 40 percent of the corpus non-English. The Hugging Face card claims 1,811 natively supported languages, including Swiss German, Romansh, and a long tail that Llama, Qwen, and Gemma largely treat as an afterthought. If your workload is English-only reasoning, that breadth buys you little. If it is not, this is one of the few sub-10B options built for it.
Architecturally it stays conventional in the ways that matter for inference: decoder-only transformer, dense, no mixture-of-experts routing. Every one of the 8B parameters is active on every token, so memory footprint and throughput scale predictably with quantization level. The release also ships in two checkpoints: Apertus-8B-2509 (base) and Apertus-8B-Instruct-2509 (chat-tuned). For anything conversational, pull the instruct variant.
Dense is the right call at this size. There is no router to add latency or to make quantized inference behave unpredictably, and the KV cache grows linearly and predictably with context. The practical consequence is that a long-context session on an 8B dense model is memory-bound, not compute-bound: you can push context further than you would on a larger model, but you pay for it in VRAM.
The fully-open training stack matters for a narrower audience. If you need to audit what data went in, verify opt-out compliance, or fork the recipe for a domain-specific run, Apertus gives you the artifacts to do it. Almost no model at this size does.
Apertus 8B is a chat model with unusually wide language coverage. Concrete workloads it fits:
Be clear about the limits. It is text-only: no vision, no audio. It is not a frontier coding model, and on English reasoning benchmarks it will trail the strongest 7B-9B competitors. Treat it as a multilingual workhorse, not a general-purpose benchmark winner.
Memory is the first constraint. Rough weight footprints by quantization:
Q4_K_M is the right starting point for most people. It costs a small amount of quality against Q8_0 and roughly halves the memory, which is the difference between "runs on a 12 GB card" and "does not." Step up to Q5_K_M or Q6_K only if you have headroom and are doing something quality-sensitive, such as translation into a low-resource language.
Hardware that realistically runs it:
Throughput ballparks for generation, single stream, Q4_K_M: roughly 100-130 tokens/sec on an RTX 4090, 45-60 on an RTX 3060 12 GB, 55-75 on an M4 Max, and 18-25 on a 16 GB M-series laptop. Context length and batch size move these numbers; benchmark your own prompt shape before committing.
Ollama is the fastest path to a working setup. It can pull GGUF builds directly from Hugging Face repos using the hf.co/ prefix, so you can point it at a community quantization of the instruct checkpoint and be chatting in a couple of minutes. For production serving, vLLM supports Apertus and will give you far better throughput under concurrency; note that it requires transformers 4.56.0 or later. llama.cpp and LM Studio both work if you prefer a local GUI or finer control over sampling.
Against Qwen3 8B: Qwen3 is the stronger generalist. It handles reasoning and code better, has a thinking mode, and ships under Apache 2.0 as well. Its language coverage is far narrower, in the low hundreds rather than low thousands. Pick Qwen3 when English and code quality dominate. Pick Apertus when the language list is the requirement.
Against Llama 3.1 8B: Llama has the larger ecosystem, more fine-tunes, and better tooling support out of the box, but it covers eight languages and ships under the Llama community license rather than Apache 2.0. If you need permissive licensing, documented training data, or non-English coverage, Apertus is the cleaner choice. If you need a drop-in model with a decade of community tooling behind it, Llama is.
The honest summary: Apertus 8B does not win on benchmark averages. It wins on language breadth, license clarity, and the fact that you can audit the entire training pipeline. Those are the reasons to choose it, and they are not small ones for teams building outside English.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every EPFL, ETH Zurich, CSCS model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.