Apertus 70B is an open-weights multilingual language model from EPFL, ETH Zurich and the Swiss National Supercomputing Centre (CSCS). It was trained on 15 trillion tokens covering more than 1,000 languages, with about 40 percent non-English data, and is intended for chatbots, translation systems and educational tools. The weights, training data and recipes are all published, and the model is released under the Apache 2.0 license. It is available on Hugging Face and through the Swiss partner Swisscom.
A solid 70B-parameter dense language model from EPFL, ETH Zurich, CSCS. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 28.7 GB | Low | |
| Q4_K_MRecommended | 43.4 GB | Good | |
| Q5_K_M | 50.4 GB | Very Good | |
| Q6_K | 58.8 GB | Excellent | |
| Q8_0 | 76.3 GB | Near Perfect | |
| FP16 | 142.8 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| NVIDIA H100 SXM5 80GBNVIDIA | SS | 62.1 tok/s | 43.4 GB |
| Google Cloud TPU v5pGoogle | SS | 51.3 tok/s | 43.4 GB |
| Intel Gaudi 2 AI AcceleratorIntel | SS | 45.5 tok/s | 43.4 GB |
| NVIDIA A100 SXM4 80GBNVIDIA | SS | 37.8 tok/s | 43.4 GB |
| Intel Gaudi 3 AI AcceleratorIntel | SS | 68.6 tok/s | 43.4 GB |
Energy cost on Corsair AI Workstation 300 (Ryzen AI Max 385) (~4.7 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Apertus 70B on Corsair AI Workstation 300 (Ryzen AI Max 385) · ~4.7 tok/s · 150W | $1.05 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 43 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA RTX A6000RunPod · Community · 48 GB VRAM | $0.33 |
NVIDIA RTX A6000RunPod · Spot · 48 GB VRAM | $0.33 |
NVIDIA A40RunPod · Community · 48 GB VRAM | $0.35 |
NVIDIA A40RunPod · Spot · 48 GB VRAM | $0.35 |
NVIDIA A40RunPod · Secure · 48 GB VRAM | $0.49 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Apertus 70B is a dense, decoder-only transformer from the Swiss AI Initiative, a collaboration between EPFL, ETH Zurich, and the Swiss National Supercomputing Centre (CSCS). It is a 70 billion parameter model trained from scratch on 15 trillion tokens, with roughly 40 percent of that data non-English, spanning more than 1,000 languages. The weights, the training data, and the full training recipes are all published, and the whole thing ships under Apache 2.0.
That combination is the reason to care. Most 70B-class open-weights models give you weights and a license that carries restrictions. Apertus gives you weights, data provenance, training code, alignment methodology, and a permissive license. For teams that need to audit what a model was trained on, or that cannot accept the legal ambiguity of a community license, this is a materially different proposition. It is also the model's main constraint: it was built to be reproducible and compliant, not to win every benchmark.
The practical framing: Apertus 70B is a general-purpose chat and translation model. It is competitive with other open models at the same 8B and 70B scale, it is unusually strong outside English, and it runs on the same hardware as any other dense 70B. If your workload is English-only code generation, other models will likely serve you better. If your workload involves European languages, low-resource languages, or any requirement that you can explain where the training data came from, Apertus is one of the few serious options.
Apertus 70B is a dense model, not a mixture-of-experts. Every one of the 70 billion parameters is active on every token. That has direct consequences for inference:
The architecture itself uses a new xIELU activation function and was trained with the AdEMAMix optimizer, both departures from the standard Llama-style recipe. Post-training combined supervised fine-tuning with alignment via QRPO. The instruction-tuned release is swiss-ai/Apertus-70B-Instruct-2509; the base model is swiss-ai/Apertus-70B.
Context length is not published as a fixed figure in the original release materials, though the model card describes long-context support. Treat the effective window as something to validate against your own workload rather than assume. The later Apertus 1.5 line explicitly extends context, so if long-context retrieval is central to your use case, check the current release before committing.
The reported pretraining data cutoff is March 2024, with some math and post-training data collected later. Plan for a knowledge horizon in early-to-mid 2024.
Apertus 70B is text-only, with chat and multilingual as its core competencies.
Multilingual chat and translation. This is where the model earns its place. The training mix is roughly 60/40 English to non-English across 1,000+ languages, and the model card claims 1,811 natively supported languages. In practice that means strong coverage of major European languages (German, French, Italian, Spanish, Dutch, Polish) and usable coverage of languages that most open models handle poorly. Concrete workloads: multilingual customer support triage, document translation pipelines where you need consistent terminology across language pairs, and localization QA.
Sovereign and regulated deployments. Because the training data is documented and the model was trained with opt-out consent respected retrospectively and with PII removal and anti-memorization measures, Apertus is designed for environments where you must answer questions about data provenance. Public sector, healthcare, and EU-regulated deployments are the obvious fits. The team publishes an EU AI Act public summary and a code of practice document.
Education and internal assistants. The stated intent covers chatbots and educational tools. A 70B dense model is a reasonable fit for a departmental assistant where you control the hardware and the data never leaves the building.
What it is not: a frontier reasoning model, a multimodal model, or a coding specialist. Do not pick it for those.
Approximate memory needed to hold weights, before KV cache and activations:
Add 4-12 GB for KV cache depending on context length and batch size. At long context, the KV cache is not a rounding error.
Use Q4_K_M unless you have a specific reason not to. It is the standard tradeoff point: roughly a quarter of the original size with quality loss that is measurable but small on chat and translation tasks. Step up to Q5_K_M or Q6_K if you have the memory headroom and care about translation fidelity on lower-resource languages, where quantization damage shows up first. Avoid Q3 and below for anything multilingual.
Rough ranges for Q4_K_M, single stream:
Batch your requests. Dense models scale throughput well with concurrency, so a server serving multiple users gets far better tokens/sec per user than a single interactive session suggests.
The model is on Hugging Face and is supported in transformers v4.56.0 and later, plus vLLM for serving. For a quick local test, Ollama is the fastest path: pull a GGUF build of Apertus-70B-Instruct and run it, or import a GGUF directly with a Modelfile if your registry does not carry it. For production serving, use vLLM with an AWQ or GPTQ 4-bit checkpoint rather than llama.cpp.
vs. Llama 3.3 70B. Same parameter class, same dense architecture, similar hardware requirements. Llama 3.3 70B is the stronger general-purpose English model and has a much larger ecosystem of fine-tunes and tooling. Apertus wins on licensing (Apache 2.0 versus the Llama Community License), on training data transparency, and on multilingual breadth. Choose Apertus when legal or provenance constraints are real; choose Llama 3.3 when raw English capability and ecosystem maturity matter more.
vs. Qwen2.5 72B. Qwen2.5 72B is also Apache 2.0 and is genuinely strong at multilingual work, particularly Asian languages, plus coding and math. It is the closest direct competitor. Apertus has an edge in European language coverage and in the completeness of its published training documentation. Qwen has an edge in benchmark performance on reasoning and code, and a larger body of community quantization work. If your language mix skews Asian, take Qwen. If it skews European, or you need the data provenance story, take Apertus.
The honest summary: Apertus 70B is not the best 70B model on a leaderboard. It is the most auditable one, and for a meaningful set of deployments that is the property that decides the purchase.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every EPFL, ETH Zurich, CSCS model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.