Mistral Large 4 is Mistral AI's flagship multimodal model, launched as a public preview on October 6, 2026. It is a mixture-of-experts model with 1.05T total parameters and 49B active parameters, accepting text and image input with a 512K-token context window. Mistral says it targets coding, agentic workflows and multimodal understanding, and plans to release the weights by the end of October 2026. Preview API pricing is $0.68 per million input tokens and $2.09 per million output tokens.
A situational 1050B-parameter MoE language model from Mistral. Pulls ahead on AA LCR (81/100), so reach for it when that's the dimension that matters. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 224.2 GB | Low | |
| Q4_K_MRecommended | 234.4 GB | Good | |
| Q5_K_M | 239.3 GB | Very Good | |
| Q6_K | 245.2 GB | Excellent | |
| Q8_0 | 257.5 GB | Near Perfect | |
| FP16 | 304.0 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| ASUS ExpertCenter Pro ET900N G3ASUS | AA | 24.4 tok/s | 234.4 GB |
| Dell Pro Max with GB300Dell | AA | 24.4 tok/s | 234.4 GB |
| HP ZGX Fury AI StationHP | AA | 24.4 tok/s | 234.4 GB |
| MSI XpertStation WS300MSI | AA | 24.4 tok/s | 234.4 GB |
| SuperMicro Super AI StationSuperMicro | AA | 24.4 tok/s | 234.4 GB |
Energy cost on AMD Instinct MI325X (~21 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Mistral Large 4 on AMD Instinct MI325X · ~21 tok/s · 1000W | $1.62 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 234 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
AMD Instinct MI350XRunPod · Community · 288 GB VRAM | $0.50 |
AMD Instinct MI350XRunPod · Spot · 288 GB VRAM | $0.50 |
AMD Instinct MI350XDigitalOcean · Spot · 288 GB VRAM | $2.46 |
AMD Instinct MI355XDigitalOcean · Spot · 288 GB VRAM | $2.97 |
AMD Instinct MI325XDigitalOcean · On-Demand · 256 GB VRAM | $3.8 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
Mistral Large 4 is Mistral AI's flagship open-weight multimodal mixture-of-experts (MoE) foundation model. Spanning 1050B total parameters with 49B active parameters per token, it represents one of the largest parameter-scale open architectures available to engineers and enterprise infrastructure teams. The model accepts interleaved text and image inputs across a massive 524,288-token context window, targeting multi-turn code synthesis, agentic orchestration, and complex document intelligence.
Positioned directly against proprietary frontier options like Claude 3.5 Sonnet and GPT-4o, Mistral Large 4 establishes a benchmark for European sovereign AI deployments. While Mistral debuted the model as a public preview API in October 2026 priced at $0.68 per million input tokens and $2.09 per million output tokens, its primary value to infrastructure engineers lies in the release of its model weights. Deploying a model of this magnitude on-premise eliminates data exfiltration risks and avoids recurring API costs for continuous batch workloads.
Operating a local AI model with 1050B parameters in 2026 demands serious hardware planning. Because of its 1.05T footprint, you cannot treat Mistral Large 4 like a typical drop-in open-source model. The MoE structure drastically reduces generation latency compared to dense models of comparable scale, but the raw memory capacity required to store the full parameter set shifts the deployment bottleneck entirely onto hardware VRAM and interconnect bandwidth.
Mistral Large 4 uses a granular mixture-of-experts design coupled with a 1.6B parameter native vision encoder. Out of the 1050B total parameters residing in memory, the routing mechanism routes each token through only 49B active parameters during any forward pass.
This routing approach defines Mistral Large 4 MoE efficiency:
The context window reaches 524,288 tokens (512K), enabling native ingestion of full codebases, deep repository trees, and hundreds of high-resolution images in a single session. However, long-context deployments heavily punish poorly planned key-value (KV) cache allocations. At full 512K sequence lengths, unquantized FP16 KV cache consumes hundreds of gigabytes independently of model weights. Serving this window locally requires paged attention mechanisms and FP8 or INT4 KV quantization in engines such as vLLM or TensorRT-LLM.
Mistral Large 4 is tuned for complex enterprise pipelines requiring precision and structural validation rather than generic conversational generation.
Benchmark evaluations reveal strong results in legal analysis, financial modeling, and cybersecurity auditing. In evaluations such as the Mistral Large 4 reasoning benchmark suite, using higher reasoning settings halves false claims during extraction tasks, though practitioners should avoid running high reasoning configurations on creative open-ended copy where factual rigidity can limit output fluidity.
Deploying this model locally requires an understanding of quantization trade-offs, host system topologies, and distributed inference backends.
To calculate storage requirements, factor in roughly 2 bytes per parameter for 16-bit precision, 1 byte for 8-bit, and 0.5 to 0.6 bytes for 4-bit variants, plus KV cache overhead.
If you are wondering how to run 1050B model on consumer GPU setups, the short answer is: you cannot run it on a single workstation card. A 24GB card like the GeForce RTX 4090 cannot even load the vision encoder and routing layers of a 1050B network, let alone the full expert layer matrix.
Even a top-spec Apple Mac Studio with an M4 Max (up to 128GB unified memory) lacks the raw capacity to load a 4-bit quantization. Running Mistral Large 4 on non-datacenter hardware requires unified memory pooling across an Apple Silicon cluster using distributed llama.cpp RPC (for instance, three networked Mac Studio M2/M3/M4 Ultra machines with 192GB unified memory each), or custom offloading to fast NVMe storage, which throttles inference down to sub-1 token per second.
The best GPU for Mistral Large 4 in production is the NVIDIA H200 (141GB SXM) or NVIDIA H100 (80GB SXM5). An 8x H200 server handles the model at FP8 precision with sufficient remaining VRAM to support multiple concurrent users running deep context requests.
For cost-managed infrastructure running 8x A100 80GB rigs, the best quantization for Mistral Large 4 is Q4_K_M or AWQ 4-bit. 4-bit quantization keeps weight memory near 600 GB, leaving around 40 GB across the cluster for inference activations and context memory.
When configured over an NVLink interconnect with an 8x H100 node:
For local testing, Ollama offers the quickest way to get started if you manage a multi-GPU environment with adequate VRAM. Pointing an Ollama or vLLM instance to a quantized GGUF or AWQ manifest initializes the pipeline with minimal manual tensor parallel configuration:
1# Example serving configuration using vLLM on an 8x 80GB node2python3 -m vllm.entrypoints.openai.api_server \3 --model mistralai/Mistral-Large-4-preview \4 --tensor-parallel-size 8 \5 --max-model-len 65536 \6 --kv-cache-dtype fp8 \7 --gpu-memory-utilization 0.95
Evaluating Mistral Large 4 against other flagship open models clarifies where its operational strengths and hardware trade-offs lie.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Mistral AI model we track.

Explore the Family
The full Mistral family leaderboard with sizes, benchmark scores, and a release timeline.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.