Atria-Dawn-Preview is a preview agentic model from Shanghai AI Laboratory, built on the 744B-parameter MoE GLM-5.2 foundation. It targets research and engineering workflows that need tool use and multi-step task completion, covering discovery, software creation, document delivery and authorized security work. It has a 256K context window, accepts text input only, and is released as open weights under an MIT license on Hugging Face and ModelScope. An FP8-quantized variant is also published.
A workable 744B-parameter MoE language model from Shanghai AI Laboratory. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 1799.5 GB | Low | |
| Q4_K_MRecommended | 1955.7 GB | Good | |
| Q5_K_M | 2030.1 GB | Very Good | |
| Q6_K | 2119.4 GB | Excellent | |
| Q8_0 | 2305.4 GB | Near Perfect | |
| FP16 | 3012.2 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | FF | 0.2 tok/s | 1955.7 GB |
| Acer Veriton GN100 AI MiniAcer | FF | 0.1 tok/s | 1955.7 GB |
| AMD Instinct MI300XAMD | FF | 2.2 tok/s | 1955.7 GB |
| AMD Instinct MI325XAMD | FF | 2.5 tok/s | 1955.7 GB |
| AMD Instinct MI355XAMD | FF | 3.3 tok/s | 1955.7 GB |
Energy cost running this model locally vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
| No featured GPU in our directory fits this model. Check back as we add more hardware. | |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 1956 GB VRAM, refreshed hourly.
No current rental listing covers this model’s VRAM requirement on the providers we track.
Atria-Dawn-Preview is a 744-billion-parameter open-weights Mixture of Experts (MoE) model developed by Shanghai AI Laboratory. Built on top of the GLM-5.2 foundation, the model is specifically engineered for autonomous research workflows, multi-step code execution, and tool-mediated digital tasks. Unlike traditional conversational models tuned purely for interactive dialogue, Atria-Dawn-Preview is structured around an agentic execution loop: formulating a hypothesis, interacting with environments via tools, verifying output state, and recovering from runtime errors.
The model occupies an aggressive position in the open-weights ecosystem. Released under a permissive MIT license with an official FP8 quantized release alongside base weights, it targets enterprise and academic teams that require full control over their infrastructure for sovereign agent workflows. Across frontier benchmarks, it is designed to compete directly with proprietary agent systems like Claude 3.5/Opus-class engines, Qwen-Max, and DeepSeek large-scale architectures, particularly in autonomous research synthesis, terminal-based software engineering, and defensive cybersecurity workflows.
Hosting a local AI model with 744B parameters in 2026 requires serious infrastructure planning. While its sparse architecture reduces the active compute footprint per forward pass, storing and addressing 744 billion parameters across high-context agent loops places heavy demands on unified memory pools, high-speed interconnects, and KV cache sizing.
Atria-Dawn-Preview operates on a sparse Mixture of Experts (MoE) backbone derived from GLM-5.2. The model maintains 744 billion total parameters distributed across specialized routing networks. While Shanghai AI Laboratory has not publicly confirmed the exact count of active parameters per token, the underlying MoE architecture activates only a sparse fraction of expert layers during any single forward pass.
This routing dynamic produces a distinct split between memory footprint and execution latency. For inference engines, the entire 744B model weight payload must reside in addressable memory (VRAM or system RAM) to avoid bus-transfer bottlenecks during routing. However, compute hardware only evaluates the active subset of experts, yielding higher Atria-Dawn-Preview MoE efficiency and faster generation throughput than a dense 744B architecture would allow.
Key architectural specifications include:
The 256K token context window is a critical component of the model's design. In autonomous workflows, agent traces accumulate rapid context overhead through tool outputs, bash command returns, compiler logs, and deep documentation retrievals. Atria-Dawn-Preview was trained using a Verifiable Experience Pipeline, exposing the network to closed-loop tool environments where state feedback and compiler error logs directly reinforce multi-step planning.
Atria-Dawn-Preview is calibrated for autonomous task resolution rather than simple single-turn query answering. Its capabilities center on four technical domains:
On complex information-gathering benchmarks such as DeepSearchQA (scoring 96.0) and BrowseComp (92.5), the model excels at formulating search queries, inspecting multi-source documents, cross-referencing claims, and organizing synthesis reports. For engineering teams, it can parse hundreds of pages of API specifications, RFC documents, or research papers and compile them into verifiable execution specs.
Using Atria-Dawn-Preview for coding scenarios extends beyond code completion. The model is trained to drive software creation from specification to passing test suites. On benchmarks like SWE-bench Pro (59.6) and Terminal-Bench 2.1 (78.3), it reads local repository directories, executes terminal commands, inspects stack traces, and iterates on bug fixes. In machine learning workflows, it can write data processing pipelines, adjust hyperparameters based on standard training output, and validate model convergence metrics.
Tool integration is central to the model's evaluation profile, registering 77.0 on the Berkeley Function-Calling Leaderboard (BFCL v4). The model adheres tightly to JSON schemas, supports nested tool calls, and handles multi-turn state dependencies where the result of one function directly shapes the arguments for the next. This makes it an ideal central controller for local agent frameworks like Codex, LangGraph, or custom agent runtimes.
On the CyberGym benchmark, Atria-Dawn-Preview achieved an 86.5 score. It handles automated vulnerability scanning, exploit reproduction, security patch validation, and regression testing within isolated containerized sandboxes.
Deploying a 744B model on local infrastructure requires a calculated balance of memory capacity, interconnect bandwidth, and quantization strategy. Practitioners looking to run Atria-Dawn-Preview locally must treat hardware sizing as an enterprise clustering or workstation problem.
The total memory required depends heavily on parameter precision and KV cache allocation for large context windows. At 256K context, the KV cache alone demands significant VRAM, especially without aggressive quantization or multi-head latent caching.
Engineers asking how to run 744B model on consumer GPU setups must adjust expectations. An individual consumer GPU like an RTX 4090 (24GB) cannot load even a fraction of this model; it would require twenty-two RTX 4090 cards running over PCIe just to load a 4-bit quantization, which suffers severe PCIe latency bottlenecks that render tokens per second unviable for agentic loops.
Viable deployment configurations include:
The best quantization for Atria-Dawn-Preview depends on latency tolerance:
For local orchestration, Ollama provides the quickest way to get started if community-converted GGUF quantizations are pulled down:
1# Running a quantized GGUF variant via Ollama (requires unified/multi-GPU mapping)2ollama run atria-dawn-preview:q4_k_m
For production server deployments, use vLLM or SGLang with tensor and pipeline parallelism across your GPU cluster:
1# Example vLLM launch across an 8-GPU node for the FP8 model2python3 -m vllm.entrypoints.openai.api_server \3 --model internlm/Atria-Dawn-Preview-FP8 \4 --tensor-parallel-size 8 \5 --max-model-len 65536 \6 --trust-remote-code
To understand Atria-Dawn-Preview's position, it must be evaluated against comparable large-scale open and semi-open MoE models, specifically DeepSeek-V3 (671B total parameters) and standard enterprise agent models.
DeepSeek-V3 is a 671B MoE model that activates 37B parameters per token. While DeepSeek-V3 is broadly trained as a generalist conversational and coding base, Atria-Dawn-Preview differentiates itself through its domain-specific post-training. Shanghai AI Laboratory tuned Atria specifically on long-horizon, tool-verified workflows:
Choose Atria-Dawn-Preview when building multi-turn autonomous agents that require deep context handling, explicit terminal execution, and verified tool sequences without cloud dependencies. If standard chat completion, pure coding assistance, or minimal hardware footprint are the primary goals, smaller MoE models or dense 70B alternatives will offer a more cost-effective local deployment path.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Shanghai AI Laboratory model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.