MiniCPM5-2B is a dense 2.5B parameter text model from OpenBMB, the second release in the MiniCPM5 series after MiniCPM5-1B. It is built for on-device and local deployment, with a 131,072 token context window and a LlamaForCausalLM architecture of 42 layers. The maker reports strengths in coding, math, long-context understanding, tool use and agentic tasks, and ships GGUF, MLX, GPTQ and LiteRT builds alongside the BF16 weights. It is released under Apache-2.0.
A solid 2.52B-parameter dense language model from openbmb. Pulls ahead on graduate-level reasoning (GPQA) (70/100), so reach for it when that's the dimension that matters. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 4.1 GB | Low | |
| Q4_K_MRecommended | 4.6 GB | Good | |
| Q5_K_M | 4.9 GB | Very Good | |
| Q6_K | 5.2 GB | Excellent | |
| Q8_0 | 5.8 GB | Near Perfect | |
| FP16 | 8.2 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Radeon RX 7600 8GBAMD | SS | 50.4 tok/s | 4.6 GB |
| NVIDIA GeForce RTX 4060NVIDIA | SS | 47.6 tok/s | 4.6 GB |
| NVIDIA GeForce RTX 5060 Ti 8GBNVIDIA | SS | 78.3 tok/s | 4.6 GB |
| AMD Radeon RX 7700 XTAMD | AA | 75.5 tok/s | 4.6 GB |
| Intel Arc B580Intel | AA | 79.7 tok/s | 4.6 GB |
Energy cost on Raspberry Pi 5 (8GB) (~5.9 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)MiniCPM5-2B on Raspberry Pi 5 (8GB) · ~5.9 tok/s · 12W | $0.067 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 5 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 2080 TiVast.ai · Spot · 11 GB VRAM | $0.04 |
NVIDIA GeForce RTX 4060Vast.ai · Spot · 8 GB VRAM | $0.04 |
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3080 TiVast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060 TiVast.ai · Spot · 8 GB VRAM | $0.05 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every openbmb model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.