Torchcast Decision 27B is a 27.8B parameter decision model from Torchcast AI, released on 2026-10-03. It answers typed questions about a state: pick one of several options, yes or no, or a rating on an ordinal scale, and returns a probability for each option. It is a LoRA fine-tune of StartLux-Decision-27B with a 262,144 token context, served through a TypeSafe-compatible /v1/systemone API. Weights are CC BY-NC 4.0, so commercial use needs permission from StartLux and Torchcast; the serving code is Apache-2.0.
A workable 27.8B-parameter dense decision model from Torchcast AI. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| Acer Veriton GN100 AI MiniAcer | SS | 74.9 GB |
| AMD Instinct MI300XAMD | SS | 74.9 GB |
| AMD Instinct MI325XAMD | SS | 74.9 GB |
| AMD Instinct MI355XAMD | SS | 74.9 GB |
| Apple M3 Ultra (32-core CPU, 80-core GPU)Apple | SS | 74.9 GB |
Cheapest current cloud rentals with at least 75 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA A100 80GB PCIeVast.ai · Spot · 80 GB VRAM | $0.22 |
NVIDIA A100 80GB SXMVast.ai · Spot · 80 GB VRAM | $0.43 |
NVIDIA A100 80GB SXMVast.ai · On-Demand · 80 GB VRAM | $0.43 |
AMD Instinct MI300XRunPod · Community · 192 GB VRAM | $0.50 |
AMD Instinct MI350XRunPod · Community · 288 GB VRAM | $0.50 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Torchcast Decision 27B is an open-weight, 27.8B parameter dense decision engine built specifically for structured probabilistic evaluation rather than standard discursive text generation. Released on October 3, 2026, by Torchcast AI, the model is a LoRA fine-tune merged into StartLux-Decision-27B, an architecture derived from Alibaba's Qwen3.8-27B foundation. Instead of returning free-form text completions, it evaluates complex system states and returns calibrated probability distributions over defined choices, binary outcomes, or ordinal ratings via a TypeSafe-compatible /v1/systemone interface.
With a native context window of 262,144 tokens, this local AI model 27.8B parameters 2026 release addresses a specific engineering bottleneck: giving autonomous agents fast, deterministic decision loops without relying on brittle prompt-engineering tricks or external API calls. On the Decision Index 0.2.1 benchmark suite, Torchcast Decision 27B scores 65.00, outperforming TypeSafe's Jev 1.13 by 7.09 points while executing single binary evaluations in 33 milliseconds on an NVIDIA H100. For engineers deploying autonomous routing, safety guardrails, or state verification pipelines locally, it provides a purpose-built alternative to general-purpose chat models.
The model is distributed under a CC-BY-NC-4.0 license, meaning self-hosted non-commercial research and local internal testing are unrestricted, but commercial production deployments require commercial agreements with both Torchcast AI and StartLux. The reference inference and evaluation codebase is licensed under Apache-2.0.
Torchcast Decision 27B operates as a dense 27.8B parameter transformer. Unlike mixture-of-experts (MoE) architectures that route active tokens across sparse parameter sub-networks, this dense design routes every token through the full parameter set. While dense execution increases compute intensity during the prefill phase, it guarantees consistent latency profiles and predictable memory bandwidth requirements across large context batches.
The model strips out the Multi-Token Prediction (MTP) head present in some base architectures, streamlining the output path to focus purely on scoring discrete options. It evaluates three specific question schemas:
choice: Selects among a discrete set of categorical possibilities.noul: Answers binary yes/no assertions.score: Produces a probability spread over an ordinal rating scale.A critical design advantage is single-pass execution. For typical requests containing multiple state questions, Torchcast Decision 27B evaluates all queries simultaneously in one forward pass rather than generating autoregressive tokens one by one.
Supporting an expansive context length of 262,144 tokens requires significant attention to memory scaling. When processing large state representations, full context processing demands FlashAttention, Flash Linear Attention, or causal-conv1d kernels to keep KV cache footprint and compute manageable. The official inference stack relies on flash-linear-attention and causal-conv1d packages to maintain execution speed.
Rather than generating descriptive text, Torchcast Decision 27B excels at decision calibration: assigning mathematically dependable probabilities to distinct options based on unstructured or structured evidence.
Torchcast Decision 27B is not designed for open-ended creative writing, conversational dialogue, or deep multi-step mathematical derivations. On deep reasoning and academic knowledge benchmarks like GPQA Diamond and MMLU-Pro, it trails knowledge-dense generalist models. It is designed to act as an analytical judge over provided context, not an encyclopedic retrieval engine.
To run Torchcast Decision 27B locally, practitioners need to plan around the sheer memory requirements of a 27.8B dense parameter network, particularly when scaling context.
In unquantized BF16 precision, the model weights alone consume 54.7 GB of storage and memory. Running native BF16 inference alongside a deep context buffer requires a minimum of 64 GB to 80 GB of VRAM.
For local practitioner rigs, 4-bit precision (such as Q4_K_M via llama.cpp or 4-bit AWQ) offers the best balance of speed, footprint, and fidelity. While extreme sub-4-bit quantizations (such as 2-bit or 3-bit) often distort probability tails and harm decision calibration, 4-bit implementations retain over 98 percent of baseline classification accuracy.
If your rig can support it, FP8 is the sweet spot for preserving calibrated probability distributions across long document inputs.
Because Torchcast Decision 27B outputs probability distributions over pre-allocated candidate tokens rather than streaming hundreds of text tokens, token generation metrics differ from standard LLMs. What matters is prefill speed and decision latency:
The quickest way to run the native server locally is through the official torchcast_decision package:
1# Clone weights and setup environment2hf download torchcast-ai/torchcast-decision-27b --local-dir torchcast-decision-27b3cd torchcast-decision-27b4python3 -m venv .venv && source .venv/bin/activate5pip install -r requirements.txt67# Verify acceleration kernels8python -m torchcast_decision.check .910# Launch local TypeSafe-compatible server11python -m torchcast_decision.server --model . --port 8090
Once running, query the endpoint directly:
1curl -s localhost:8090/v1/systemone \2 -H 'Content-Type: application/json' \3 -d '{4 "state": "System memory usage is at 94% with 400MB swap remaining.",5 "questions": [6 {"id": "q1", "type": "noul", "text": "Is memory exhaustion imminent?"}7 ]8 }'
For environments utilizing Ollama or standard GGUF tools, community conversion scripts can expose the underlying base architecture, though full compatibility with the custom /v1/systemone logit extraction heads requires running the provided Python runner.
Understanding Torchcast Decision 27B performance requires comparing it directly to relevant decision engines and similarly sized foundation models.
TypeSafe's Jev 1.13 has served as the reference implementation for typed decision serving. Torchcast Decision 27B establishes a notable lead on general decision metrics, scoring 65.00 on Decision Index 0.2.1 compared to Jev 1.13's 57.91. Torchcast leads on 31 of 38 index sub-benchmarks, notably outperforming Jev in stance evaluation, contract logic, and simulated tool calling. However, Jev 1.13 retains an advantage on broad factual retrieval and complex knowledge tasks (leading Torchcast by 5.8 points in the Knowledge index area). Choose Torchcast for agent control flow and local automation; stick with Jev if your decision points rely heavily on open-domain world knowledge.
StartLux-Decision-27B is the direct architectural parent of Torchcast Decision 27B. While StartLux achieves a respectable 63.88 on Decision Index 0.2.1, Torchcast's LoRA fine-tuning and recalibrated head refine probability distributions in edge cases, reducing probability overconfidence by 1.12 index points overall. If choosing between the two for new local deployments, Torchcast Decision 27B offers better probability calibration across diverse prompts.
Deploying a standard open-weight model (such as a generic Qwen or Gemma chat variant) for classification tasks requires structured JSON prompting or logit masking wrappers. These approaches are prone to schema drift, require continuous generation of multi-token reasoning chains to achieve high accuracy, and consume significant compute. Torchcast Decision 27B provides native classification logits in a single forward pass, delivering predictable latency and calibrated output probabilities without conversational overhead.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Torchcast AI model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.