Beam is Reflection AI's first open-weight model, a sparse mixture-of-experts text model with 501B total parameters and 23B active. It targets coding, reasoning and agentic workloads, with a 1 million token context window and pretraining on 23.8 trillion tokens. Reflection says it is competitive with larger open models such as GLM 5.2 while using three to four times less inference compute, though those claims have not been independently verified. Weights, a model card and a technical report are expected later in October 2026 under an Apache 2.0 license.
A workable 501B-parameter MoE language model from Reflection AI. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 193.0 GB | Low | |
| Q4_K_MRecommended | 197.8 GB | Good | |
| Q5_K_M | 200.1 GB | Very Good | |
| Q6_K | 202.9 GB | Excellent | |
| Q8_0 | 208.7 GB | Near Perfect | |
| FP16 | 230.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| AMD Instinct MI355XAMD | SS | 32.6 tok/s | 197.8 GB |
| AMD Instinct MI325XAMD | AA | 24.4 tok/s | 197.8 GB |
| ASUS ExpertCenter Pro ET900N G3ASUS | AA | 28.9 tok/s | 197.8 GB |
| Dell Pro Max with GB300Dell | AA | 28.9 tok/s | 197.8 GB |
| HP ZGX Fury AI StationHP | AA | 28.9 tok/s | 197.8 GB |
Energy cost on Framework Desktop (AMD Ryzen AI Max+ PRO 495, 192GB) (~1.1 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Beam on Framework Desktop (AMD Ryzen AI Max+ PRO 495, 192GB) · ~1.1 tok/s · 120W | $3.60 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 198 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
AMD Instinct MI350XRunPod · Community · 288 GB VRAM | $0.50 |
AMD Instinct MI350XRunPod · Spot · 288 GB VRAM | $0.50 |
AMD Instinct MI350XDigitalOcean · Spot · 288 GB VRAM | $2.46 |
AMD Instinct MI355XDigitalOcean · Spot · 288 GB VRAM | $2.97 |
AMD Instinct MI325XDigitalOcean · On-Demand · 256 GB VRAM | $3.8 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Beam is an open-weight foundation model developed by Reflection AI, featuring a sparse Mixture-of-Experts (MoE) architecture with 501 billion total parameters and 23 billion active parameters per token. Released under the permissive Apache 2.0 license, it is designed for demanding engineering, automated reasoning, and tool-augmented agentic execution. Pretrained on 23.8 trillion curated tokens and refined through an extensive reinforcement learning campaign, Beam represents an effort to bring frontier-tier performance into self-hosted enterprise environments without the proportional compute penalties of traditional ultra-dense networks.
Positioned directly against major open-weight models such as GLM 5.2 and Qwen 3.8-Max, Beam prioritizes inference compute efficiency over brute parameter activation. Reflection claims the model matches GLM 5.2 on advanced reasoning while requiring three to four times less inference compute per generated token. For organizations seeking a local AI model with 501B parameters in 2026 that can be audited, modified, and hosted on private infrastructure, Beam offers an alternative to closed APIs, provided the hosting hardware can satisfy its substantial memory footprint.
Its 1,000,000-token context window allows developers to ingest entire software repositories, complex legal documentation, or multi-turn agent logs directly into working memory. While the model delivers the execution speed of a 23B parameter network during generation, running it on local hardware requires infrastructure capable of storing all 501 billion weights in memory.
Beam relies on a sparse Mixture-of-Experts design. In a standard dense model, every parameter participates in calculating every token. In Beam, a gating router evaluates each token and routes it through only 23 billion of the total 501 billion parameters.
This decoupling of total capacity from active computation creates the core advantage of Beam MoE efficiency:
The architectural tradeoff is straightforward: inference throughput is fast because the compute path processes only 23B parameters, but the static hardware requirement remains defined by the 501B total parameter weight matrix.
Beam is a text-only model engineered specifically for technical workflows, software engineering, and multi-step autonomous tool use.
Reflection built Beam for coding across complex environments rather than isolated code completion. On agentic coding benchmarks, Beam posts an 80.9 score on SWE-bench Verified and 78.0 on SWE-bench Multilingual. The model excels at:
In synthetic reasoning, formal proofs, and complex logic benchmarks, Beam matches GLM 5.2. Its high-compute reinforcement learning phase prioritized chain-of-thought verification across sandboxed execution tests, reducing hallucination rates when calculating numeric problems, parsing legal contracts, or designing system architectures.
The combination of a 1M token context window and structured schema adherence allows Beam to orchestrate multi-tool agentic workflows. It can parse extensive API specifications, select appropriate endpoints, generate verified JSON payloads, and maintain coherence across long-running autonomous tasks without degradation.
Deploying Beam on private infrastructure eliminates API rate limits and keeps proprietary code completely on-premises. However, hosting a 501B parameter architecture requires planning around memory capacity and interconnect speeds.
To run Beam locally, your hardware must hold the entire 501B parameter set in memory, even though only 23B are active at any given moment.
| Quantization Level | Weight Size (VRAM) | Minimum Recommended VRAM (with 32K context) | High-Context Headroom (1M context) | Recommended Use Case |
|---|---|---|---|---|
| FP16 / BF16 | ~1,002 GB | 1,100 GB+ | 1,250 GB+ | Enterprise multi-node data centers |
| Q8_0 | ~530 GB | 580 GB | 720 GB | Minimal precision loss, multi-GPU rigs |
| Q4_K_M | ~290 GB | 330 GB | 450 GB | Best quantization for Beam (optimal efficiency) |
| Q3_K_M | ~220 GB | 250 GB | 370 GB | Tighter budgets, minor reasoning loss |
| Q2_K | ~160 GB | 190 GB | 300 GB | Emergency memory fit, noticeable quality loss |
For most practitioners, Q4_K_M represents the sweet spot, preserving reasoning fidelity while lowering weight size to approximately 290 GB.
Engineers frequently ask how to run a 501B model on consumer GPU setups. A standard consumer card like the Nvidia GeForce RTX 4090 (24 GB VRAM) cannot run Beam on its own. Even a dedicated rig of four RTX 4090s provides only 96 GB of VRAM, which falls well short of the ~290 GB needed for a 4-bit quant.
Practical deployment requires workstation or enterprise-grade memory architecture:
Q4_K_M at standard context lengths. An 8x A100/H100 node is recommended if you plan to saturate the 1,000,000-token context window.Q4_K_M weights directly. While generation speeds on unified RAM will be slower than dedicated GDDR6X or HBM3 setups, it represents a viable local alternative to a full enterprise server rack.Once loaded into memory, Beam tokens per second are dictated by its 23B active parameter pathway rather than its 501B total size. On a 4x H100 node running vLLM or TensorRT-LLM, you can expect generation speeds between 30 and 45 tokens per second for standard context tasks. On Apple Silicon unified memory, memory bandwidth constraints typically throttle throughput to around 6 to 12 tokens per second.
For rapid local testing, inference engines like vLLM or optimized builds of llama.cpp serve as the most straightforward runtimes. For containerized deployments, Ollama provides an accessible path once community GGUF quantizations are downloaded to an appropriately provisioned machine.
Evaluating Beam against comparable open-weight models highlights its structural efficiency and specific operating compromises.
GLM 5.2 is a widely deployed open-weight alternative with high marks in raw coding and technical reasoning.
Qwen 3.8-Max targets raw benchmark dominance at enterprise scale.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Reflection AI model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.