AREX-2 is a 27B dense multimodal model from BAAI (Beijing Academy of Artificial Intelligence) built for long-horizon agent work. It improves a candidate solution over multiple test-time rounds by proposing, measuring, reflecting and revising, and it is trained on machine-learning and algorithmic-programming tasks with verifiable feedback plus AREX deep-research data. Context length is 262,144 tokens, weights are open on Hugging Face, and the license is Apache 2.0. Reported scores include 70.7 on Frontier-CS, 81.8 on MLE-Lite, 84.0 on BrowseComp and 92.2 on GAIA.
A solid 27.36B-parameter dense language model from BAAI. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 68.0 GB | Low | |
| Q4_K_MRecommended | 73.7 GB | Good | |
| Q5_K_M | 76.5 GB | Very Good | |
| Q6_K | 79.8 GB | Excellent | |
| Q8_0 | 86.6 GB | Near Perfect | |
| FP16 | 112.6 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| Intel Gaudi 3 AI AcceleratorIntel | SS | 40.4 tok/s | 73.7 GB |
| NVIDIA H200 SXM 141GBNVIDIA | SS | 52.4 tok/s | 73.7 GB |
| AMD Instinct MI300XAMD | SS | 57.9 tok/s | 73.7 GB |
| Google TPU v7 (Ironwood)Google | SS | 80.6 tok/s | 73.7 GB |
| NVIDIA B200 GPUNVIDIA | SS | 87.3 tok/s | 73.7 GB |
Energy cost on NVIDIA A100 SXM4 80GB (~22 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)AREX-2 on NVIDIA A100 SXM4 80GB · ~22 tok/s · 400W | $0.599 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 74 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA A100 80GB SXMVast.ai · Spot · 80 GB VRAM | $0.13 |
NVIDIA A100 80GB PCIeVast.ai · Spot · 80 GB VRAM | $0.40 |
NVIDIA A100 80GB SXMVast.ai · On-Demand · 80 GB VRAM | $0.47 |
AMD Instinct MI300XRunPod · Community · 192 GB VRAM | $0.50 |
AMD Instinct MI300XRunPod · Spot · 192 GB VRAM | $0.50 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
AREX-2 is a 27.36B dense multimodal model from BAAI (Beijing Academy of Artificial Intelligence), the lab behind the original AREX deep-research agent family. It occupies a specific niche: long-horizon agent work where the model improves a candidate solution across multiple test-time rounds instead of answering once and stopping. The loop is propose, measure, reflect, revise. Training combines machine-learning and algorithmic-programming tasks with verifiable feedback plus AREX deep-research data, and the self-improvement behavior learned on those tasks transfers to research workloads without adding new search trajectories.
The numbers that matter for evaluation: 70.7 on Frontier-CS, 81.8 on MLE-Lite, 84.0 on BrowseComp, and 92.2 on GAIA. Those are competitive with far larger open-weight models on agentic and ML-engineering benchmarks, and they were produced by a dense 27B checkpoint, not a mixture-of-experts with a large activated-parameter budget.
Why this matters for local deployment: 27.36B dense fits on a single 24GB consumer card at 4-bit quantization, and the weights are open on Hugging Face under Apache 2.0. That combination (agentic tool use, vision input, 262,144-token context, permissive license, single-GPU footprint) is rare in this parameter class. If you are building an agent harness that needs to run on hardware you own, AREX-2 is one of the few serious candidates.
AREX-2 is a dense transformer, not a mixture-of-experts. Every one of the 27.36B parameters is active on every token. Two practical consequences follow:
It is built on a Qwen3.8-27B backbone, so it is a Qwen3.8-compatible multimodal model: a vision encoder feeding a text decoder. The Hugging Face pipeline tag is image-text-to-text, and weights ship as safetensors for transformers.
The 262,144-token context window is not a marketing figure here, it is load-bearing. The reflective loop depends on the model reading its own prior state: scores, logs, stack traces, error messages, timings, and the accumulated record of what it already tried. Long context is what lets a single agent session hold an entire ML experiment history or a multi-hour research transcript without an external summarizer.
Declared capabilities are chat, code, vision, reasoning, and function-calling. The realistic use cases follow directly from the training mix.
Where it is not the right pick: high-volume short-answer chat, broad multilingual work, or anything that needs a generalist's breadth. AREX-2 is specialized.
Weights only, before KV cache:
KV cache is the second budget. At 262,144 tokens the cache is very large (tens of GB depending on attention configuration). For local work, cap context at 32K-64K, which costs a few GB and covers most agent sessions. Full 262K context realistically requires 48GB or more of memory.
llama.cpp, but expect single-digit tokens per second. Not viable for agent loops.Q4_K_M is the right starting point for most users. Move to Q5_K_M or Q6_K if you have memory headroom, particularly for ML engineering tasks where numeric reasoning and log parsing matter. Q8_0 is worth it only on 32GB+ hardware. Below Q4, the reflective loop degrades first: the model starts misreading its own logs and error output, which defeats the entire point of the architecture.
On an RTX 4090 at Q4_K_M, expect roughly 20-35 tokens/sec on short generations, dropping as context fills. An M4 Max lands nearer 10-20 tokens/sec. Prompt processing is the real bottleneck for long agent transcripts, so prefill speed matters more than raw decode speed for this workload.
Ollama is the fastest path. Check the Ollama library for an AREX-2 tag; if none exists yet, pull a GGUF build from Hugging Face or convert the safetensors yourself with llama.cpp's convert_hf_to_gguf.py. Note that vision requires the matching mmproj file, and not every GGUF release ships one. For multi-user serving or high-throughput batch work, run the native weights under vLLM instead.
AREX-2 vs Qwen3-32B. Qwen3-32B is a dense generalist with strong coding and reasoning and a large ecosystem of quants and tooling. It has no vision in the base checkpoint and no long-horizon self-improvement training. Pick Qwen3-32B for broad chat, multilingual work, and mature local tooling. Pick AREX-2 when the task is iterative with verifiable feedback and you want vision plus 262K context in the same checkpoint.
AREX-2 vs Gemma 3 27B. Gemma 3 27B is dense, multimodal, and runs at a similar memory footprint, but with a 128K context window and a generalist post-training mix. It is the better choice for straightforward vision Q&A and general assistant work. AREX-2 wins on agentic benchmarks, ML engineering, and any workload where the model must sustain productive iteration across many rounds.
The tradeoff is consistent: AREX-2 gives up some generalist breadth for specialization in long-horizon, feedback-driven tasks. If your workload looks like an experiment loop or a research agent, that trade is worth making. If it looks like a chatbot, it is not.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every BAAI model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.