Blink v0.3 26B-A4B is a decision model from Pixilab.ai, built as a full fine-tune of google/gemma-4-26B-A4B-it. It reads a state, a question and a list of options and returns a probability for every option in one forward pass, for routing, moderation, tagging, gating, dedup and yes/no checks. It has 25.8B total parameters with 4B active, a 32k context window, and ships under Apache 2.0. This NVFP4 build takes 17.5 GB on disk and runs on a single RTX PRO 5000.
A solid 25.8B-parameter MoE decision model from Pixilab.ai. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 3.9 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 3.9 GB |
| AMD Instinct MI300XAMD | SS | 3.9 GB |
| AMD Instinct MI325XAMD | SS | 3.9 GB |
| AMD Instinct MI355XAMD | SS | 3.9 GB |
Cheapest current cloud rentals with at least 4 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3060Vast.ai · Spot · 12 GB VRAM | $0.05 |
NVIDIA GeForce RTX 3060Vast.ai · On-Demand · 12 GB VRAM | $0.06 |
NVIDIA GeForce RTX 5060 TiVast.ai · Spot · 16 GB VRAM | $0.07 |
NVIDIA Tesla V100 16GBVast.ai · Spot · 16 GB VRAM | $0.09 |
NVIDIA RTX A4000Vast.ai · Spot · 16 GB VRAM | $0.09 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Blink v0.3 26B-A4B is a specialized decision model developed by Pixilab.ai, designed to replace generic conversational LLMs in structured evaluation, classification, and routing tasks. Built as a full fine-tune of Google's google/gemma-4-26B-A4B-it, the model abandons conversational prose in favor of deterministic, single-token probability evaluation. When supplied with a system state, an inquiry, and a set of candidate options, Blink scores every option in a single forward pass, returning calibrated probabilities directly from the model logprobs.
Deploying large autoregressive models simply to generate "yes" or "no" tokens or structured JSON payloads introduces unnecessary latency, parsing brittleness, and compute overhead. Blink v0.3 26B-A4B addresses this bottleneck by functioning as a dedicated "System 1" decision engine. It delivers competitive classification accuracy without requiring chain-of-thought generation, making it an essential local AI model with 25.8B parameters for production infrastructure in 2026.
With 25.8B total parameters and a sparse Mixture-of-Experts (MoE) design that activates only 4B parameters per token, Blink delivers enterprise-tier routing capabilities while remaining lightweight enough to run on single-GPU workstations. The release is licensed under Apache 2.0, permitting unrestricted commercial and private deployment.
Blink v0.3 is built upon the Gemma 4 MoE architecture (Gemma4ForConditionalGeneration). Its total parameter footprint is 25.8B, but its dynamic routing mechanism limits execution to 4B active parameters per forward pass.
Understanding the Blink v0.3 26B-A4B MoE efficiency is critical for hardware sizing:
Blink conforms to the Surogate Decisions v1 protocol. The model does not use speculative decoding or reasoning scratchpads. Its internal "thinking" mode must be kept disabled, as generating multi-step reasoning tokens degrades its calibration and decision latency. Instead, candidate choices are mapped to letters (A, B, C, etc.), and calibrated probabilities are extracted via a softmax over the logprobs of the first generated token position. The model ships pre-calibrated for an inference temperature of 1.3, achieving an Expected Calibration Error (ECE) of just 0.016 over more than 216,000 scored benchmarks.
Blink v0.3 26B-A4B is purpose-built to act as a programmable switch inside agentic frameworks and high-throughput pipelines. Key capabilities include:
For yes/no operations, Blink uses a standardized criterion where option "A" maps to "No" and option "B" maps to "Yes". For multi-class decision trees involving more than 26 options, the protocol seamlessly handles double-letter identifiers (AA, AB, AC).
To run Blink v0.3 26B-A4B locally, practitioners need to plan for memory capacity rather than compute throughput. Because inference terminates after one generated token, traditional generation speed is less relevant than time-to-first-token (TTFT) and prompt ingestion speed.
The best quantization for Blink v0.3 26B-A4B depends on your target accelerator architecture:
When selecting the best GPU for Blink v0.3 26B-A4B:
Because Blink only activates 4B parameters during execution, prompt evaluation speeds regularly exceed 150 to 300 tokens per second on an RTX 4090, depending on context length. Generating the single evaluation token occurs virtually instantaneously (often sub-10ms). The overall Blink v0.3 26B-A4B tokens per second metric is dominated by prompt-processing speed rather than autoregressive decoding loops.
The standard production deployment uses vLLM to serve the model as an OpenAI-compatible endpoint with logprob extraction enabled:
1vllm serve PixilabAI/Blink-v0.3-26B-A4B-NVFP4 \2 --served-model-name blink \3 --quantization modelopt_fp4 \4 --kv-cache-dtype fp8 \5 --max-model-len 32768 \6 --enable-prefix-caching \7 --chat-template-content-format string
Once running, client requests should set max_tokens=1, temperature=0, logprobs=True, and top_logprobs=20. Pass enable_thinking: false in the template kwargs to prevent the model from entering non-deterministic conversational loops.
For fast local testing, Ollama provides an alternative path using GGUF builds:
1ollama run blink-26b-a4b:q4_k_m
Note that when using standard chat wrappers like Ollama, you must query raw logprobs via the backend API to access calibrated decision probabilities rather than just parsing the output string.
Evaluating Blink v0.3 26B-A4B vs Surogate Rune 26B-A4B v3 shows how specialized decision models have converged. On the 38-benchmark Decision Index 0.2.1 suite, Blink v0.3 posts an aggregate score of 57.48, marginally edging out Surogate Rune v3 (57.44). Blink demonstrates clear superiority in language nuance (64.3 vs 63.1) and artistic/editorial judgement (43.8 vs 41.9), while Rune v3 holds a slight lead on tool routing (71.2 vs 70.0).
Compared to larger dense deciders like Decider Chat Gemma-4-31B (Index score 57.33), Blink offers equivalent decision fidelity while requiring significantly less compute during the forward pass due to its 4B active parameter profile. However, dense 31B alternatives still hold an advantage in deep encyclopedic knowledge and complex multi-hop deductive puzzles (such as GPQA Diamond, where Blink v0.3 lands at 27.2).
If your pipeline requires conversational generation, continuous prose output, or raw code writing, deploy a standard chat model like Gemma 4 or Qwen 2.5. If your architecture demands an ultra-fast, local, deterministic routing and evaluation checkpoint that eliminates fragile regex string parsing, Blink v0.3 26B-A4B is currently one of the most cost-efficient open-weight choices available.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Pixilab.ai model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.