IQuest-Q1 is a sparse Mixture-of-Experts model from IQuest, built for agentic coding, reasoning and multi-step tool use. It has about 320B total parameters with roughly 15B activated per token, 88 layers and 256 experts, and supports a 524,288 token context. Weights are published openly on Hugging Face under an "other" license, and deployment is documented for SGLang and vLLM with OpenAI-compatible serving.
A solid 320.32B-parameter MoE language model from IQuest. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 69.0 GB | Low | |
| Q4_K_MRecommended | 72.1 GB | Good | |
| Q5_K_M | 73.6 GB | Very Good | |
| Q6_K | 75.4 GB | Excellent | |
| Q8_0 | 79.2 GB | Near Perfect | |
| FP16 | 93.4 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| Intel Gaudi 3 AI AcceleratorIntel | SS | 41.3 tok/s | 72.1 GB |
| NVIDIA H200 SXM 141GBNVIDIA | SS | 53.6 tok/s | 72.1 GB |
| AMD Instinct MI300XAMD | SS | 59.2 tok/s | 72.1 GB |
| Google TPU v7 (Ironwood)Google | SS | 82.4 tok/s | 72.1 GB |
| NVIDIA B200 GPUNVIDIA | SS | 89.3 tok/s | 72.1 GB |
Energy cost on NVIDIA A100 SXM4 80GB (~23 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)IQuest-Q1 on NVIDIA A100 SXM4 80GB · ~23 tok/s · 400W | $0.586 |
GPT-6.1 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 72 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA A100 80GB SXMVast.ai · Spot · 80 GB VRAM | $0.13 |
NVIDIA A100 80GB PCIeVast.ai · Spot · 80 GB VRAM | $0.40 |
NVIDIA A100 80GB SXMVast.ai · On-Demand · 80 GB VRAM | $0.47 |
AMD Instinct MI300XRunPod · Community · 192 GB VRAM | $0.50 |
AMD Instinct MI300XRunPod · Spot · 192 GB VRAM | $0.50 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Cumulative cost of buying the smallest device in our directory that runs this model at 4-bit, compared with paying flagship APIs for the same token volume.
Assumes 1 million tokens a day, a 70/30 input and output mix and $0.12 per kWh. The device only runs as long as the workload needs. Use the ROI calculator for your own numbers.
IQuest-Q1 is a sparse Mixture-of-Experts language model from IQuest Research, released with open weights on Hugging Face. It carries 320.32B total parameters but activates only about 15B per token, a sparse ratio of roughly 21:1. That ratio is the whole point of the model: you get the representational capacity of a 320B-class network while paying inference costs closer to a 15B one.
The model is aimed at a specific job, agentic coding and multi-step tool use. It is not a general-purpose chat model with coding bolted on. IQuest trained it around software engineering, long-horizon task execution, and tool calling, with reinforcement learning and multi-teacher on-policy distillation layered on top of the base pretrain. It reads and writes English and Chinese, handles function-calling natively, and supports a 524,288-token context window.
That combination puts it in direct competition with Qwen3-235B-A22B, GLM-4.5 (355B total, 32B active), and DeepSeek's V-series MoE releases. IQuest's own benchmark chart positions it against DeepSeek's latest V4 releases on agentic coding, terminal operation, and security tasks. The practical question for anyone reading this page is not whether it scores well on a leaderboard, but whether a 320B MoE is deployable on hardware you actually have. For most individuals, it is not. For teams with a multi-GPU node, it is.
IQuest-Q1 uses a decoder-only Transformer backbone with sparse MoE layers. The configuration:
Two design choices matter for anyone sizing hardware.
First, active parameters govern both speed and working-set memory. Each token routes through 8 of 256 experts plus the shared components, so weight reads during decode are roughly 15B parameters worth per step, not 320B. The catch is that all 320.32B must still be resident somewhere: VRAM, system RAM, or disk. On a multi-GPU server this is fine. On a workstation it is the binding constraint.
Second, the hybrid attention pattern keeps KV cache survivable at long context. With three of every four layers using 4,096-token sliding windows, only 22 of the 88 layers attend to the full sequence. In BF16, that works out to roughly 47 GB of KV cache at the full 524,288 tokens, against what would be closer to 190 GB with uniform full attention. Budget for that separately from weights.
The multi-token prediction head is worth noting for throughput: at inference, a single recursive MTP layer with depth 8 lets the model draft multiple tokens per forward pass, which is where a substantial part of its decode speed advantage comes from.
IQuest-Q1 is built for work that runs for minutes to hours, not single-turn queries.
iquest_q1) and a matching reasoning parser for SGLang, and it is documented against harnesses like Claude Code and Codex.It has no vision or audio path. If your pipeline injects images into the agent conversation, IQuest's own evaluation notes say multimodal content is replaced with placeholders during tokenization. Plan for a separate model if you need grounding on screenshots or diagrams.
Weight storage at the major precisions, excluding KV cache and runtime overhead:
| Precision | Weights (approx.) | Minimum practical GPU setup |
|---|---|---|
| BF16 | ~640 GB | 8x H200 141GB (1128 GB total) |
| FP8 / INT8 | ~320 GB | 8x H100 80GB, or 4x H200 141GB |
| Q4_K_M | ~180 GB | 3x H100 80GB, or 2x H200 141GB |
| Q3_K_M | ~145 GB | 2x H100 80GB |
| Q2_K | ~110 GB | 2x A100 80GB (tight) |
Add roughly 47 GB of KV cache in BF16 for the full 524,288-token context, or about half that with FP8 KV cache. Shorter contexts cut this proportionally.
Be realistic: no single consumer GPU runs this. An RTX 4090 with 24 GB, an RTX 5090 with 32 GB, or an M4 Max with up to 128 GB cannot hold even a Q2_K build. The only consumer-adjacent path is a large-memory unified system such as a Mac Studio with 256 GB or more, or a workstation with 256 GB of system RAM plus GPU offload through llama.cpp or KTransformers. Expect single-digit to low-double-digit tokens per second in those configurations, and expect the MoE routing to spend most of its time on memory bandwidth rather than compute.
For a single-socket desktop with a 24 GB card, plan on a smaller MoE in the 30B-active class instead. The 15B active budget is cheap to compute, but 320.32B of weights have to live somewhere.
There are no official GGUF builds listed at launch, so quantization is a community and post-release question rather than a documented path. Where GGUF or AWQ builds appear, Q4_K_M is the sensible default: roughly 4.5 bits per weight, about 180 GB of storage, and minimal quality loss on structured tasks like code editing. Q3_K_M saves about 35 GB at a noticeable cost on reasoning depth. Avoid Q2 builds for agentic work; the failure mode is subtler tool-call syntax errors that are hard to debug.
For multi-GPU serving, FP8 is usually the better trade. It halves the footprint versus BF16, keeps numerics well within tolerance for this architecture, and preserves the tensor-parallel kernels that SGLang and vLLM are tuned for.
IQuest documents two production paths, both OpenAI-compatible.
SGLang with an 8-way tensor parallel split:
1python -u -m sglang.launch_server \2 --model-path "$MODEL_ROOT" \3 --served-model-name IQuest-Q1 \4 --tp-size 8 \5 --dtype bfloat16 \6 --attention-backend fa3 \7 --mem-fraction-static 0.85 \8 --tool-call-parser iquest_q1 \9 --reasoning-parser iquest_q1
A prebuilt image (iquestlabworkspace/sglang-iquest-q1:cu130) and a fork of vLLM maintained by IQuest are both published. Recommended sampling for reproducibility is temperature=1.0, top_p=0.95, top_k=20.
Ollama is the fastest way to get most models running, but it is not the right entry point here. Unless and until official quantized GGUF artifacts appear in the Ollama library, treat ollama pull as unavailable and start from SGLang or vLLM.
With 15B active parameters and MTP depth 8, decode is fast for a model of this size. On 8x H100 80GB at FP8, expect roughly 60-120 tokens/sec single-stream and substantially higher aggregate at batch. On a 2x H200 Q4 setup, expect roughly 20-40 tokens/sec single-stream. These figures move a lot with context length, batch size, and KV cache precision, so benchmark your own workload before committing to a hardware purchase.
vs. Qwen3-235B-A22B. Qwen3 activates 22B per token against 235B total, so it is cheaper per token and fits on fewer GPUs. IQuest-Q1 activates 15B, which is lower, but its 320.32B total makes it heavier to host. If your bottleneck is GPUs you already own, Qwen3-235B-A22B is the easier fit and has a broader ecosystem of quants. If you are optimizing for per-token cost at scale and can afford the footprint, Q1's lower active count and 524,288-token context are the stronger proposition for long agent runs.
vs. GLM-4.5 (355B total, 32B active). GLM-4.5 is a generalist with strong tool-calling and a mature serving story. IQuest-Q1 is narrower and more coding and terminal-task oriented, with a context window nearly an order of magnitude longer. Choose Q1 when your workload is repository-level software engineering or long-horizon agent execution. Choose GLM-4.5 when you need broad general capability from a single deployment.
vs. DeepSeek's V-series MoE. DeepSeek's larger MoE models remain the reference point for open reasoning quality, and IQuest benchmarks against them directly. They also activate more parameters per token and generally require comparable or larger hardware. Q1's case is efficiency per activated parameter, not raw ceiling.
One licensing note before you commit: the weights are released under an other license specific to IQuest-Q1, not Apache 2.0 or MIT. Read the LICENSE file in the Hugging Face repository before shipping anything commercial.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every IQuest model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.