Naive-N0.5-Flash is an open-weight Mixture-of-Experts model from NaiveAI, released September 27, 2026, and built for coding and AI research and development. It has 309B total parameters with 15.5B active and a native 1M-token context window using a hybrid of sliding-window attention and DeepSeek Sparse Attention, with no full-attention layers. Weights and inference code are under the MIT license, and API pricing is listed at $0.10 per million input tokens and $0.40 per million output tokens. The model needs FP8-capable NVIDIA GPUs and about 315 GB for weights.
A workable 309B-parameter MoE language model from NaiveAI. A pragmatic middle-ground choice when you need open weights without a flagship-sized footprint. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
See how different quantization levels affect VRAM requirements and quality for this model.
| Format | VRAM Required | Quality | |
|---|---|---|---|
| Q2_K | 130.2 GB | Low | |
| Q4_K_MRecommended | 133.5 GB | Good | |
| Q5_K_M | 135.0 GB | Very Good | |
| Q6_K | 136.9 GB | Excellent | |
| Q8_0 | 140.8 GB | Near Perfect | |
| FP16 | 155.5 GB | Full |
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | Speed | VRAM |
|---|---|---|---|
| Google TPU v7 (Ironwood)Google | SS | 44.5 tok/s | 133.5 GB |
| NVIDIA B200 GPUNVIDIA | SS | 48.2 tok/s | 133.5 GB |
| AMD Instinct MI355XAMD | SS | 48.2 tok/s | 133.5 GB |
| AMD Instinct MI325XAMD | SS | 36.2 tok/s | 133.5 GB |
| AMD Instinct MI300XAMD | SS | 32.0 tok/s | 133.5 GB |
Energy cost on Apple M4 Max (40-core GPU) (~3.3 tok/s, Q4_K_M) vs flagship API pricing.
| Source | Cost per 1M tokens |
|---|---|
Local (energy only)Naive-N0.5-Flash on Apple M4 Max (40-core GPU) · ~3.3 tok/s · 92W | $0.931 |
GPT-6 SolOpenAI · in $2.00 · out $10.00 | $4.40 |
Claude Opus 5.5Anthropic · in $4.00 · out $20.00 | $8.80 |
Gemini 3.5 FlashGoogle · in $1.50 · out $9.00 | $3.75 |
Grok 4.5xAI · in $2.00 · out $6.00 | $3.20 |
API prices blended at 70% input / 30% output.
Hardware amortisation not included. Run the full ROI calculator for payback math.
Cheapest current cloud rentals with at least 133 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
AMD Instinct MI300XRunPod · Community · 192 GB VRAM | $0.50 |
AMD Instinct MI300XRunPod · Spot · 192 GB VRAM | $0.50 |
NVIDIA H200 NVLRunPod · Community · 141 GB VRAM | $0.50 |
NVIDIA H200 NVLRunPod · Spot · 141 GB VRAM | $0.50 |
NVIDIA H200 SXMVast.ai · Spot · 141 GB VRAM | $1.32 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Naive-N0.5-Flash is a 309B-parameter Mixture-of-Experts model from NaiveAI, released September 27, 2026, under the MIT license. It activates 15.5B parameters per token and carries a native 1,000,000-token context window. Weights and inference code are open. That combination, frontier-scale sparse capacity plus a permissive license, is the reason this model matters to anyone building locally.
The category it occupies is narrow and deliberate: coding and AI research and development. NaiveAI trained it to participate in its own R&D loop, writing code, running experiments, and iterating on results, with human researchers setting direction. That shows up in the benchmark suite the company reports (SWE-Bench Pro, Terminal-Bench 2.1, ALE-CLI, MLE-bench-30, PaperBench), which skews heavily toward agentic software engineering and research tasks rather than chat.
The headline architectural claim is that there are no full-attention layers anywhere in the 48-layer stack. NaiveAI replaced global attention with a hybrid of Sliding-Window Attention (SWA) and DeepSeek Sparse Attention (DSA), betting that a fully local-or-sparse network can hold coherence at 1M tokens. If it works, it removes the quadratic decoding cost that makes million-token inference impractical on most hardware. It is the first frontier-scale open-weight release to attempt this.
The numbers that govern local deployment:
Active parameters are the number that determines compute per token, not VRAM. At 15.5B active, the FLOP cost of a forward pass is closer to a mid-size dense model than to a 309B one. That is the MoE efficiency story, and it is why NaiveRT can quote 50 tokens/s per user in Standard mode and up to 2,000 tokens/s in Ultrafast mode on NaiveAI's own infrastructure.
VRAM is a different problem. Every expert has to be resident somewhere, or fetched over PCIe at a cost that destroys throughput. The 309B total is what you have to fit, not the 15.5B.
The attention design is what makes 1M context tractable. Most layers use a 128-token sliding window with per-token decoding cost that does not grow with sequence length. The nine DSA layers replace global attention: a lightweight indexer scores the full history, and the backbone attends only to the top 2,048 selected tokens. The full KV cache is still retained and the indexer still scans the whole history, so this reduces attention compute and memory traffic rather than eliminating long-context cost entirely. NaiveAI adapted the model to this structure through continued pretraining: 50B tokens of indexer warmup, 3T tokens of sparse attention training, and 200B tokens of learning rate decay, 3.25T tokens total.
This is a text-only model. No vision, no audio. Its stated capabilities are chat, code, and reasoning, with the emphasis clearly on the first two-thirds of that list being applied to software.
Concrete workloads where it fits:
It is not the right pick for latency-sensitive single-turn chat, for multimodal work, or for anything requiring a small memory footprint.
Weight size is the binding constraint. Approximate figures by precision:
| Precision | Approx. weight size | Notes |
|---|---|---|
| FP8 (native) | ~315 GB | Requires FP8-capable NVIDIA GPUs |
| Q6_K | ~230 GB | Near-lossless |
| Q4_K_M | ~160 GB | The practical default |
| Q3_K_M | ~120 GB | Noticeable degradation on reasoning |
| Q2_K | ~90 GB | Last resort |
Add KV cache on top. Even with GQA4 and sparse attention, the full KV cache is retained, and at 1M tokens it runs to tens of gigabytes. Budget headroom above the weight figure.
A single RTX 4090 (24 GB) cannot run this model at any usable quantization. Neither can an M4 Max at 128 GB, except at Q2 with heavy CPU offload and token rates that make it a science experiment rather than a tool.
Realistic configurations:
Expect roughly 20-40 tokens/s decode on an 8x4090 Q4 setup, higher on H100 clusters. The 2,000 tokens/s Ultrafast figure is NaiveAI's serving stack on their hardware (mega-kernel fusion, Programmatic Dependent Launch, speculative decoding), not a number you should expect from a workstation.
Q4_K_M is the right starting point for most local users. It cuts weights to roughly half of FP8 with modest quality loss and is the only precision that fits comfortably on 8x24GB cards. Move to Q6_K if you have 384 GB and care about reasoning fidelity. Avoid Q3 and below for coding tasks: the degradation shows up in exactly the multi-step agentic behavior this model is built for.
Ollama is the fastest path if a GGUF build exists for your target precision. Check that first, because the hybrid SWA-DSA stack is nonstandard: there are no full-attention layers, so generic attention kernels do not apply, and llama.cpp support depends on kernels for both the sliding-window and sparse paths. If GGUF support is not there, use NaiveAI's released inference code (MIT licensed) or a serving framework with a compatible implementation. Do not assume a stock vLLM build will load it without checking the architecture registry.
Against DeepSeek-V4.1-Flash, the closest open-weight competitor at this scale, the tradeoff is attention design. DeepSeek's sparse attention work is the direct ancestor of the DSA layers here, but Naive-N0.5-Flash pushes further by removing full attention entirely and pairing it with a 5:1 SWA-DSA layout. If your workload lives at 500K-1M tokens, that is the differentiator. If you run mostly short-context coding, the advantage narrows and you should compare on benchmark fit for your specific task.
Against Kimi-K3, another frontier-scale open-weight MoE, the decision usually comes down to hardware and license. Naive-N0.5-Flash is MIT, which is the most permissive option in this weight class and removes legal friction for commercial deployment. Kimi-K3's tooling ecosystem is more mature in some serving stacks.
Choose Naive-N0.5-Flash when you need a permissive license, million-token context, and strong agentic coding at a compute cost closer to a 15B dense model. Look elsewhere when you need multimodal input, a sub-100 GB footprint, or a validated serving path on a single consumer GPU.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every NaiveAI model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.