Gemma is Google’s open-weight model family, derived from the same research as Gemini. Sizes range from 270M on-device variants to 27B server-class checkpoints, with strong fine-tuning support.
See all models from GoogleModels in family
8
Open weight
8
API only
0
Avg score
59.6
Top benchmark
89.1
Arena Score
Total HF downloads
32.5M
Primary modality
Text
Context window
128K – 256K
First release
Mar 2025
Latest release
Jun 2026
Every release in the Gemma family, ranked by composite score across benchmarks, popularity, efficiency, and versatility.
| # | Model | Modality | Score | Params | Released |
|---|---|---|---|---|---|
| 1 | text | AA71.2 | 12B | — | |
| 2 | text | BB68.7 | 26B | Apr 2026 | |
| 3 | text | BB64.1 | 12B | Jun 2026 | |
| 4 | text | BB59.9 | 31B | Apr 2026 | |
| 5 | text | BB58.4 | 4B | Apr 2026 | |
| 6 | text | BB56.2 | 4B | Mar 2025 | |
| 7 | text | CC50.7 | 2B | Apr 2026 | |
| 8 | text | CC47.7 | 27B | Mar 2025 |
When each release shipped, newest first. Useful for tracking version cadence.
Jun 3
Apr 1
Apr 1
Apr 1
Apr 1
Mar 11
Mar 11
Composite grades across this family. Higher is better, blending benchmarks, popularity, and efficiency.
Models with downloadable weights, ranked by composite score.
| # | Model | Modality | Score | Params | Released |
|---|---|---|---|---|---|
| 1 | text | AA71.2 | 12B | — | |
| 2 | text | BB68.7 | 26B | Apr 2026 | |
| 3 | text | BB64.1 | 12B | Jun 2026 | |
| 4 | text | BB59.9 | 31B | Apr 2026 | |
| 5 | text | BB58.4 | 4B | Apr 2026 | |
| 6 | text | BB56.2 | 4B | Mar 2025 | |
| 7 | text | CC50.7 | 2B | Apr 2026 | |
| 8 | text | CC47.7 | 27B | Mar 2025 |
The Gemma family is a series of AI models from Google. This page lists every release in the family with its benchmark scores, parameter count, and hardware requirements.
By composite score, Gemma 4 12B Coder is currently the top model in the family. For local inference, match the parameter count to your VRAM budget. For quality, pick the highest scorer that fits.
See the open-weight section above for models you can run locally. The API-only section lists closed releases that must be accessed through the provider’s API.
Spin up an instance in the cloud, or pick local hardware that fits.
Advertising disclosure: we earn commissions when you shop through the links below.
Vast.ai
Decentralized GPU marketplace with the lowest hourly prices.
RunPod
Pay-per-second GPU rentals starting at $0.20/hr.
Digital Ocean
Spin up a GPU droplet in minutes, starting at $0.75/hr.
Vultr
GPU cloud with hourly and monthly plans, starting at $0.50/hr.
GPU Mart
Dedicated GPU servers and VPS billed monthly, starting at $0.50/hr.
PPQ.ai
Multi-model inference gateway for production workloads.
Find Local Hardware
See which GPU, Mac, or workstation can run Gemma on-prem.
Local LLM Mini PCs
Compact machines that run Gemma on your desk. Chosen for local inference.