Perplexity Decider v1.1 27B is an open-weights decision model from Perplexity, released on October 6, 2026 as a full fine-tune of Qwen3.8-27B. It returns probabilities for yes/no, choice and rubric-score questions instead of text, and scores 61.56 on the Hugging Face Decision Index 0.3, up from 56.4 for v1. Weights are Apache-2.0 and self-hosting needs about 49 GiB of GPU memory plus Perplexity's bundled inference code, since default causal inference does not reproduce the evaluated setup. The hosted API costs $0.02 per million input tokens via POST /v1/decisions.
A workable 27.8B-parameter dense decision model from Perplexity. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| Acer Veriton GN100 AI MiniAcer | SS | 18.4 GB |
| AMD Instinct MI300XAMD | SS | 18.4 GB |
| AMD Instinct MI325XAMD | SS | 18.4 GB |
| AMD Instinct MI355XAMD | SS | 18.4 GB |
| Apple M3 Ultra (32-core CPU, 80-core GPU)Apple | SS | 18.4 GB |
Cheapest current cloud rentals with at least 18 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3090Vast.ai · Spot · 24 GB VRAM | $0.12 |
NVIDIA RTX PRO 4000 BlackwellVast.ai · Spot · 24 GB VRAM | $0.12 |
NVIDIA GeForce RTX 3090Vast.ai · On-Demand · 24 GB VRAM | $0.14 |
NVIDIA RTX A5000RunPod · Community · 24 GB VRAM | $0.16 |
NVIDIA RTX 4000 AdaRunPod · Community · 20 GB VRAM | $0.20 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Perplexity Decider v1.1 27B is an open-weights dense decision model built by Perplexity, released under the Apache-2.0 license. Rather than operating as a conventional autoregressive text generator that outputs tokens sequentially, Decider v1.1 acts as a scoring and classification engine. It evaluates input state against structured criteria, returning calibrated probability distributions for binary questions, categorical choices, and discrete rubric scores.
Built as a full fine-tune of Qwen3.8-27B, the model achieves an overall score of 61.56 on the Hugging Face Decision Index 0.3 benchmark, surpassing both its predecessor (v1 at 56.4) and comparable architectures like TypeSafe's Jev (57.9). By replacing open-ended text generation with direct logit readouts, Decider v1.1 eliminates parsing errors, JSON schema violations, and output non-determinism in mission-critical routing and evaluation pipelines.
For engineers seeking to run Perplexity Decider v1.1 27B locally, this checkpoint presents distinct architectural traits compared to standard causal models. Local deployment requires accounting for custom readout heads, noncausal attention masks, and substantial VRAM footprints during inference.
Perplexity Decider v1.1 27B uses a dense 27.8B parameter backbone based on Qwen3.8-27B. Unlike standard decoder-only large language models that map hidden states to a massive vocabulary matrix, Decider swaps out the standard language model head (lm_head) for a specialized readout module.
Key architectural characteristics include:
readout.safetensors head with dimensions [255, 5120] in BF16 format. This head extracts classification logits directly from the final transformer layer.pipeline("text-generation") cannot reproduce the evaluated results out of the box. Running the model requires either using Perplexity's bundled autojev Python package or exporting the readout rows into specific vocabulary token slots with an engine patched for bidirectional attention.Decider v1.1 27B is optimized for high-throughput decision-making where deterministic, numerical outputs are mandatory. Because it does not generate text, it cannot draft responses, write code, or engage in conversation. Instead, it provides high-precision evaluations across three core primitives:
In multi-agent architectures and complex support systems, routing user intent to the correct agent or API tool typically requires parsing a generative LLM's response. Decider v1.1 handles categorical selection natively. A single forward pass yields exact probabilities across predefined categories (such as triage between billing, technical bug reports, or account security), allowing application logic to route traffic based on raw floating-point thresholds.
For automated safety policies, Decider processes inputs against binary policy criteria. Instead of asking a model to output "SAFE" or "UNSAFE" and parsing string tokens, Decider returns a calibrated probability of violation. If the probability exceeds an operational risk tolerance (such as 0.85 for mild violations or 0.50 for severe risks), the system triggers intervention instantly.
Evaluating complex artifacts, such as grading user-written essays, checking code against stylistic constraints, or assessing customer support transcripts, usually requires structured output parsing. Decider natively evaluates input against multi-tiered rubric criteria, computing probabilities across discrete point scales and providing a statistically sound expected score.
Deploying this local AI model with 27.8B parameters in 2026 requires understanding its memory footprint and execution path. Because Decider uses a forward evaluation pass rather than continuous token decoding, hardware demands differ significantly from standard generative LLMs.
In its native BF16 format, the model weights take approximately 49 GiB of storage. Inference requires additional working memory for the bidirectional attention computation across the 8,192-token context window.
| Precision / Quantization | Memory Required (Weights) | Recommended GPU VRAM | Target Hardware Setup |
|---|---|---|---|
| BF16 (Native) | ~49 GiB | 56 GiB - 64 GiB | 1x NVIDIA A100/H100 (80GB), 2x RTX 4090 (48GB combined via NVLink/PCIe), Apple M3/M4 Max (64GB+) |
| INT8 / Q8_0 | ~28 GiB | 32 GiB - 36 GiB | 1x RTX 6000 Ada (48GB), 2x RTX 3090/4090, Apple Mac Studio (48GB+) |
| INT4 / Q4_K_M | ~16 GiB | 20 GiB - 24 GiB | 1x NVIDIA RTX 4090 (24GB), 1x RTX 3090 (24GB), Apple M4 Pro (24GB+) |
For practitioners running on consumer silicon, Q4_K_M (or 4-bit AWQ/EXL2 equivalent) is the best quantization for Perplexity Decider v1.1 27B. At 4-bit precision, the parameter weight footprint drops to around 16 GiB. This fits comfortably inside the 24GB VRAM buffer of an NVIDIA RTX 4090 or RTX 3090, leaving roughly 6 GiB to 8 GiB for the KV cache and working activation buffers at the maximum 8,192 context limit.
INT8 (Q8_0) offers near-zero degradation from BF16, but requires at least 32 GiB of addressable VRAM, which forces users onto multi-GPU systems or professional workstation cards.
Evaluating Perplexity Decider v1.1 27B tokens per second requires a different mindset than testing conversational chatbots:
Standard inference runners like Ollama require custom configuration for this model. Because Ollama assumes a standard autoregressive model with an lm_head, loading Decider directly produces invalid results unless the checkpoint is exported into an expanded vocabulary mapping and paired with an engine that supports lifting the causal attention mask.
For reliable production serving, deploy using the official bundled runtime:
perplexity-ai/pplx-decider-v1.1-27b).source/src from the repository to your Python path.DecisionModel class pointing to the local weight directory on a CUDA or MPS device.state text and the question schema directly to model.predict().Understanding where Decider v1.1 fits requires comparing it against its dedicated architectural predecessor, TypeSafe's Jev, and conventional generative models like Qwen 2.5 32B.
| Feature / Metric | Perplexity Decider v1.1 27B | TypeSafe Jev | Qwen 2.5 32B (Generative) |
|---|---|---|---|
| Architecture | Dense 27.8B (Bidirectional) | Dense 27B (Bidirectional) | Dense 32.5B (Causal Autoregressive) |
| Output Type | Direct Probabilities / Readout | Direct Probabilities / Readout | Text / Structured JSON Tokens |
| Decision Index 0.3 | 61.56 | 57.90 | N/A (Generative proxy ~52-55) |
| Latency per Decision | Sub-200ms (Single pass) | Sub-200ms (Single pass) | 800ms - 2,500ms (Token generation) |
| VRAM (Native BF16) | ~49 GiB | ~48 GiB | ~65 GiB |
| Context Length | 8,192 tokens | 8,192 tokens | 128,000 tokens |
| License | Apache-2.0 | Proprietary / Apache | Apache-2.0 |
Both models utilize custom readout heads to bypass generative tokenization. However, Decider v1.1 scores 3.66 points higher on the Decision Index 0.3 benchmark overall. Decider v1.1 demonstrates substantial advantages in the Language (69.45 vs. 62.0) and Retrieval (61.26 vs. 55.4) subcategories, largely due to training on broader task-oriented corpora like tasksource and optimizing noncausal layer transitions. Jev holds a narrow edge in raw Knowledge retrieval (51.4 vs. 48.18), but Decider v1.1 is significantly more reliable across programmatic classification workflows.
A general-purpose generative LLM like Qwen 2.5 32B can perform classification by prompting it to output a JSON object or a single label. However, doing so incurs severe operational tradeoffs:
For engineers building agentic control planes, routing fabrics, and data labeling pipelines, Perplexity Decider v1.1 27B provides a faster, mathematically sound alternative to prompting generative language models.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Perplexity model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.