Decision-2.0-Vega-27B is the 27B model in the Decision 2.0 family from the vLLM Semantic Router team, released on 2 October 2026. It takes text or JSON input plus a set of questions and answers choice, yes/no and score questions in one forward pass, returning a probability for every option without generating text. It is a LoRA adapter on Qwen3.8-27B with 29.37B parameters and a 32,768 token context, licensed Apache-2.0. The model card reports a Jev Decision Index of 56.5 and a median latency of 71.4 ms per single-question request on one GPU.
A workable 29.37B-parameter dense decision model from vLLM Semantic Router. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| Acer Veriton GN100 AI MiniAcer | SS | 25.2 GB |
| AMD Instinct MI300XAMD | SS | 25.2 GB |
| AMD Instinct MI325XAMD | SS | 25.2 GB |
| AMD Instinct MI355XAMD | SS | 25.2 GB |
| Apple M3 Ultra (32-core CPU, 80-core GPU)Apple | SS | 25.2 GB |
Cheapest current cloud rentals with at least 25 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA A100 80GB PCIeVast.ai · Spot · 80 GB VRAM | $0.22 |
NVIDIA RTX PRO 5000 BlackwellVast.ai · Spot · 48 GB VRAM | $0.25 |
NVIDIA GeForce RTX 5090Vast.ai · Spot · 32 GB VRAM | $0.30 |
NVIDIA A40Vast.ai · Spot · 48 GB VRAM | $0.31 |
NVIDIA RTX A6000RunPod · Community · 48 GB VRAM | $0.33 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Decision-2.0-Vega-27B is a 29.37B parameter dense model developed by the vLLM Semantic Router team, designed specifically for zero-generation classification, semantic routing, and structural scoring. Released under the Apache-2.0 license, Vega-27B is constructed as a trained LoRA adapter on top of Qwen3.8-27B. Instead of using autoregressive decoding to output structured text, JSON blobs, or tool-call tags, the model evaluates input data against multiple predefined questions simultaneously in a single forward pass, outputting raw probability distributions across all choices without generating text tokens.
This single-pass mechanism addresses a critical bottleneck in agentic and automated inference systems: routing latency. When orchestrating multi-agent workflows, utilizing a standard generative LLM to categorize intents, extract boolean assertions, or assign triage ratings consumes substantial compute cycles and introduces parsing fragility. Vega-27B eliminates text generation overhead entirely, clocking a median single-GPU latency of 71.4 ms for single-question requests while scoring a class-leading 74.0 on JevArena and 56.5 on the Jev Decision Index (0.2.1 kit).
For local deployments, Vega-27B functions as a high-capacity System-1 routing engine. It acts as an authoritative front-line filter capable of evaluating user state against complex routing trees, leaving heavier, latency-intensive generative models free to execute actual task fulfillment.
Decision-2.0-Vega-27B is built on a dense transformer architecture totaling 29.37B parameters, using Qwen3.8-27B as its foundation base. Rather than generating response sequences token by token, the model couples its underlying representation layers with specialized decision heads via parameter-efficient fine-tuning (PEFT).
The model ingests two primary inputs: an arbitrary state (plain text or structured JSON) and a dictionary of structured questions. In a single forward pass, the model processes the context and computes normalized log-probabilities across every candidate option.
Key technical specifications include:
Because Vega-27B does not generate sequences autoregressively, its operational profile differs sharply from standard text models. Traditional LLM inference requires maintaining and expanding an autoregressive KV cache across dozens or hundreds of generated output steps. Vega-27B runs strictly in prefill mode: the input prompt and question definitions are processed together, the output logits are mapped directly to the defined answer slots, and the process completes immediately. This keeps memory footprints deterministic and drastically reduces execution variance.
Vega-27B is engineered for fast, programmatic infrastructure decisions where reliability and low latency are non-negotiable. Its native decision heads support three primary modalities:
These capabilities enable several concrete deployment architectures:
In multi-agent systems, incoming user queries must be dispatched to specialized agents or toolchains. Vega-27B handles complex intent routing across dozens of candidate endpoints in under 100 ms, supplying calibrated confidence scores for each route. If the top probability falls below an engineering threshold, the system can immediately trigger an escalation fallback.
Instead of running heavy moderation prompts through an external API, you can run Decision-2.0-Vega-27B locally to evaluate inputs against strict compliance policies. A single forward pass can simultaneously check for prompt injection patterns, policy violations, and PII presence across the full 32,768-token window.
Structured data pipelines often suffer from JSON syntax errors or hallucinated keys when using generative models for classification. Because Vega-27B outputs log-probabilities directly over predefined answer sets, output parsing failures are mathematically impossible.
Deploying a local AI model with 29.37B parameters in 2026 requires understanding its memory profile. Because Vega-27B does not generate text, you do not need to budget VRAM for dynamic autoregressive generation tokens, though processing large contexts up to 32k tokens still requires adequate working memory.
The memory required to load the model depends on the precision level:
Searches regarding Decision-2.0-Vega-27B tokens per second reflect a misunderstanding of how the model functions. Because the engine does not output generative text tokens, traditional metrics like generation tokens per second do not apply.
Instead, Decision-2.0-Vega-27B performance is measured in batch forward-pass latency and queries per second (QPS):
Loading the model requires Python with transformers >= 5.17, PyTorch, Safetensors, and PEFT. Because the custom decision heads and system_one execution path are bundled with the repository, you must enable trust_remote_code=True.
1import json2from transformers import AutoModel34model = AutoModel.from_pretrained(5 "vllm-sr/Decision-2.0-Vega-27B",6 trust_remote_code=True,7 device_map="auto"8)910result = model.system_one(11 state="Customer received broken display unit. Demanding refund immediately.",12 questions={13 "department": {14 "type": "choice",15 "instructions": "Route to correct department",16 "criteria": {17 "support": "General hardware defects and returns",18 "billing": "Invoices and credit card disputes",19 "sales": "New orders and product inquiries"20 }21 },22 "escalation_required": {23 "type": "yes_no",24 "instructions": "Does the request require manager escalation?"25 },26 "severity": {27 "type": "score",28 "instructions": "Rate urgency",29 "criteria": ["Low", "Medium", "Critical"]30 }31 }32)3334print(json.dumps(result["answers"], indent=2))
For environments where Ollama or standard GGUF runtimes are preferred for local execution, community GGUF conversions and ONNX runtimes are available, though running via native Transformers or vLLM preserves direct access to the model.system_one() method.
When evaluating Decision-2.0-Vega-27B vs AutoJev-27B, Vega-27B holds the top benchmark ranking among identical-size decision models:
| Metric / Model | Decision-2.0-Vega-27B | AutoJev-27B | Eikos-27B | Jebadiah-27B |
|---|---|---|---|---|
| Parameters | 29.37B | 27B class | 27B class | 27B class |
| JevArena Score | 74.0 | 72.1 | 69.3 | 65.5 |
| Human-Labelled Transfer | 58.7 | 58.7 | 58.7 | 57.8 |
| Jev Decision Index | 56.5 | Not reported | Not reported | Not reported |
| Execution Style | Single-pass (PEFT) | Single-pass | Single-pass | Single-pass |
Engineers frequently attempt routing tasks using general-purpose models like Qwen2.5-32B-Instruct or Mistral-Small-24B paired with JSON mode or function calling.
If your local architecture requires generating freeform text, Vega-27B is the wrong choice. If your pipeline demands deterministic categorization, dynamic routing, or safety screening at hardware line-rate, Decision-2.0-Vega-27B provides an optimal balance of 29.37B classification capacity and sub-100ms local execution.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every vLLM Semantic Router model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.