Quasar 438B is a 438-billion-parameter reasoning model from Multiverse Computing, built for enterprise agents and coding. It is a mixture-of-experts model derived from Z.ai's GLM-5.2, with its expert layers compressed from 265 to 148 experts per layer using the company's CompactifAI pipeline. It carries a 1 million token context window, runs in English and Spanish, and scores 43 on the Artificial Analysis Intelligence Index v4.1.1. It is served through the CompactifAI API.
A situational 438B-parameter MoE language model from Multiverse Computing. Pulls ahead on AA LCR (76/100), so reach for it when that's the dimension that matters.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Average benchmark score against active parameters for every text model we track. Models higher up deliver more quality for their size.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 3666.6 GB | Low | |
| Q4_K_MRecommended | 3758.5 GB | Good | |
| Q5_K_M | 3802.3 GB | Very Good | |
| Q6_K | 3854.9 GB | Excellent | |
| Q8_0 | 3964.4 GB | Near Perfect | |
| FP16 | 4380.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | FF | 0.1 tok/s | 3758.5 GB |
| Acer Veriton GN100 AI MiniAcer | FF | 0.1 tok/s | 3758.5 GB |
| AMD Instinct MI300XAMD | FF | 1.1 tok/s | 3758.5 GB |
| AMD Instinct MI325XAMD | FF | 1.3 tok/s | 3758.5 GB |
| AMD Instinct MI355XAMD | FF | 1.7 tok/s | 3758.5 GB |
Energy cost running this model locally vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
| No featured GPU in our directory fits this model. Check back as we add more hardware. | |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Sonnet 5.5Anthropic · in $2.00 · out $10.00 | $4.40 |
Gemini 4 ArgonGoogle · in $2.00 · out $10.00 | $4.40 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 3759 GB VRAM, refreshed hourly.
No current rental listing covers this model’s VRAM requirement on the providers we track.
Quasar 438B is a 438-billion-parameter mixture-of-experts reasoning model from Multiverse Computing. It is built for enterprise agents and coding, not general-purpose chat. The model is derived from Z.ai's GLM-5.2 and compressed with Multiverse's CompactifAI pipeline, which reduced its expert layers from 265 to 148 experts per layer. It carries a 1 million token context window, runs in English and Spanish, and scores 43 on the Artificial Analysis Intelligence Index v4.1.1. That score makes it the highest-scoring European model on that index, ahead of Mistral Medium 3.5 (30), NVIDIA Nemotron 3 Ultra (38), and Inkling (42). Claude Opus 5 leads at 63.
Multiverse Computing is a Spanish company focused on model compression and efficient AI. Quasar is their flagship model, served through the CompactifAI API. There is no public weight release, no license specified, and no training cutoff published. The active parameter count for the MoE architecture is also undisclosed. That matters for local deployment: total parameters set VRAM requirements, active parameters set compute per token.
What Quasar does well is specific: multi-step reasoning, tool use, code execution, and large-context tasks. It returns 500 tokens, including thinking time, in 15.3 seconds. Only a few models in its comparison set are faster, and most that outscore it take far longer. It is a specialist, tuned for agentic and coding work.
Quasar 438B uses a mixture-of-experts architecture with 438B total parameters. Multiverse Computing has not disclosed the number of active parameters per token. That number determines inference speed and compute cost, so without it, tokens-per-second estimates remain approximate. The expert layers were pruned from 265 to 148 experts per layer using a quantum-inspired method from the CompactifAI pipeline. Instead of scoring experts individually, the method treats expert selection as a joint problem across the layer. This preserves capabilities that matter as a set for agentic and coding tasks. A healing pass then recovers and sharpens behavior in those domains.
The model was compressed with quantization-aware techniques into FP8 and NVFP4 variants. The current CompactifAI API serves the FP8 version. The model supports an OpenAI-compatible chat endpoint, function and tool calling, structured outputs via response_format, and reasoning. Reasoning is always on and cannot be toggled off. You can set reasoning_effort to high or max (default is max).
The 1 million token context window is the headline spec for practitioners. It enables whole-repository code analysis, long agent traces, legal or financial document review, and multi-hour conversation memory. The model is text-only. There is no vision or audio input.
Quasar 438B is tuned for coding and agentic workflows. Concrete uses include multi-file code refactoring, test generation, terminal-based automation, and code review over an entire repository. It scores on evaluations like Terminal-Bench v2.1 and SciCode, which measure code execution and scientific coding. For enterprise agents, it supports planning, tool use, function calling, and structured outputs. That makes it suitable for workflow automation, API orchestration, and customer support agents that need to call external services.
Multilingual support covers English and Spanish. That is narrower than most frontier models, but useful for bilingual enterprise deployments, Spanish-language customer support, and cross-border workflows. Reasoning is always on, with high and max effort levels. This suits complex analysis, compliance checks, and financial modeling where multi-step logic matters.
It is not a general-purpose model. It is not for image generation, audio, casual chat, or broad world knowledge outside its tuned domains. Multiverse Computing describes it as a specialist, and the pruning objective targeted agentic and coding tasks. Capability outside that domain is where the reduction was taken.
As of publication, Multiverse Computing has not published public weights or a license for Quasar 438B. There is no GGUF release and no Ollama package listed. The model is served through the CompactifAI API. So "run Quasar 438B locally" today means running a client against that API, not loading weights into your GPU. The fastest way to test it is the CompactifAI API, which uses an OpenAI-compatible endpoint.
If weights become available, here is the hardware math based on 438B total parameters:
No single consumer GPU can run it. An RTX 4090 has 24 GB. An M4 Max tops out at 128 GB unified memory. Neither can hold even a Q4 version. Multi-GPU consumer setups like 4x RTX 4090 (96 GB) are also insufficient. The best GPU for Quasar 438B, if weights were available, would be an H100 80GB, RTX 6000 Ada 48GB, or Mac Studio M3 Ultra 512GB for Q4.
Expected tokens per second depends on active parameter count, which is undisclosed. For comparable 400B+ MoE models with 30B-50B active parameters, expect 20-60 tok/s on multi-GPU H100 systems and 5-15 tok/s on Mac Studio with offloading. Without the active count, those are estimates, not guarantees.
For quantization, Q4_K_M is the practical choice for most users. It cuts memory roughly in half versus FP8 with modest quality loss. Because the model was already pruned for agentic and coding tasks, Q4 is usually safe. Use FP8 if you have 500 GB+ VRAM and need maximum accuracy. NVFP4 is for Blackwell GPUs. Q8 is for accuracy-critical work with ample VRAM.
Ollama is not an option today. If Multiverse Computing releases GGUF weights, ollama run quasar-438b would be the quickest way to get started. Check the CompactifAI docs for current access.
Against NVIDIA Nemotron 3 Ultra, Quasar is smaller (438B vs 550B total) but scores higher on the Artificial Analysis Intelligence Index (43 vs 38). Nemotron may offer broader general capability and different licensing. Choose Quasar for agentic coding, long context, and European deployment. Choose Nemotron if you need a larger model with a different ecosystem or open weights.
Against Mistral Medium 3.5, Quasar scores 43 vs 30. Mistral is a European model with broader language support. Choose Quasar for reasoning-heavy agentic tasks, 1M token context, and coding. Choose Mistral for general chat, broader multilingual work, and lighter hardware.
Against GLM-5.2, Quasar is the compressed, domain-tuned derivative. GLM-5.2 offers open weights and broader general performance. Choose GLM-5.2 if you want local weights and general capability. Choose Quasar if you want the specialized agentic and coding version with 1M context, FP8 serving, and API access. For local deployment today, GLM-5.2 is the more realistic option because Quasar's weights are not public.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Multiverse Computing model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.