Muse Glimmer-30B is a dense open-weight model from Meta Superintelligence Labs, released in August 2026 under an Apache 2.0 license. It is built for always-on local agent workflows such as sequential tool calling, local coding, failure recovery and LLM-as-a-judge evaluation on a Mac or a single consumer GPU. It has 30 billion parameters, a 131,072 token context window, and a 1.8B parameter perception encoder that accepts interleaved text and images. Weights are published on Hugging Face with integrations for llama.cpp, MLX, ExecuTorch and vLLM.
A solid 30B-parameter dense language model from Meta. Pulls ahead on competition math (AIME 2026) (95/100), so reach for it when that's the dimension that matters.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 43.1 GB | Low | |
| Q4_K_MRecommended | 49.4 GB | Good | |
| Q5_K_M | 52.4 GB | Very Good | |
| Q6_K | 56.0 GB | Excellent | |
| Q8_0 | 63.5 GB | Near Perfect | |
| FP16 | 92.0 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| Google Cloud TPU v5pGoogle | SS | 45.1 tok/s | 49.4 GB |
| Intel Gaudi 2 AI AcceleratorIntel | SS | 40.0 tok/s | 49.4 GB |
| NVIDIA H100 SXM5 80GBNVIDIA | SS | 54.6 tok/s | 49.4 GB |
| Intel Gaudi 3 AI AcceleratorIntel | SS | 60.3 tok/s | 49.4 GB |
| NVIDIA H200 SXM 141GBNVIDIA | SS | 78.3 tok/s | 49.4 GB |
Energy cost on Corsair AI Workstation 300 (Ryzen AI Max 385) (~4.2 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Muse Glimmer-30B on Corsair AI Workstation 300 (Ryzen AI Max 385) · ~4.2 tok/s · 150W | $1.20 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 49 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA A100 80GB PCIeVast.ai · Spot · 80 GB VRAM | $0.24 |
NVIDIA A100 80GB SXMVast.ai · Spot · 80 GB VRAM | $0.32 |
AMD Instinct MI300XRunPod · Community · 192 GB VRAM | $0.50 |
AMD Instinct MI300XRunPod · Spot · 192 GB VRAM | $0.50 |
NVIDIA H200 NVLRunPod · Community · 141 GB VRAM | $0.50 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Meta model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.