Strands Decider 2B is an open source decision model from Amazon Web Services' Strands Labs, released October 1, 2026. It is built from the Qwen3.5-2B LLM torso with a LoRA adapter and a small readout head, so instead of generating text it answers typed questions: yes/no (noul), one of N options (choice), or a level on an ordered scale (score), each with a calibrated confidence. At 2 billion parameters it runs on a local CPU or GPU and returns answers in tens of milliseconds. Weights, training data and scripts are published on Hugging Face and GitHub under Apache-2.0.
A solid 2B-parameter dense decision model from Amazon Web Services. Treat the modality benchmarks above as the leading indicator of fit — composite scoring across modalities is still maturing. Newly released, so production-readiness is still being shaken out.
Generated from this model’s benchmarks and ranking signals. Editor reviews refine it over time.
Access model weights, configuration files, and documentation.
No benchmark data available for this model yet.
The top devices for this model at 4-bit, ranked by fit and speed.
| Device | Grade | VRAM |
|---|---|---|
| ACEMAGIC M1A Pro (i9-13900HK + ARC A770)ACEMAGIC | SS | 1.7 GB |
| Acer Veriton GN100 AI MiniAcer | SS | 1.7 GB |
| AMD Instinct MI300XAMD | SS | 1.7 GB |
| AMD Instinct MI325XAMD | SS | 1.7 GB |
| AMD Instinct MI355XAMD | SS | 1.7 GB |
Cheapest current cloud rentals with at least 2 GB VRAM, refreshed hourly.
Advertising disclosure: we earn commissions when you shop through the links below.
| Option | Cost / GPU-hour |
|---|---|
NVIDIA GeForce RTX 3070RunPod · Community · 8 GB VRAM | $0.13 |
NVIDIA GeForce RTX 3070RunPod · Spot · 8 GB VRAM | $0.13 |
NVIDIA RTX A5000RunPod · Community · 24 GB VRAM | $0.16 |
NVIDIA RTX A5000RunPod · Spot · 24 GB VRAM | $0.16 |
NVIDIA GeForce RTX 3080RunPod · Community · 10 GB VRAM | $0.17 |
Per-GPU rate across RunPod, the Vast.ai marketplace, DigitalOcean, and Vultr.
Spot tier is interruptible. Plan for restarts when comparing against on-demand prices.
Strands Decider 2B is an open source decision model from Amazon Web Services' Strands Labs, released October 1, 2026. It is not an LLM. It is built from the Qwen3.5-2B LLM torso with a LoRA adapter and a small readout head, so instead of generating text it answers typed questions: yes/no (noul), one of N options (choice), or a level on an ordered scale (score), each with a calibrated confidence. At 2 billion parameters it runs on a local CPU or GPU and returns answers in tens of milliseconds. Weights, training data and scripts are published on Hugging Face and GitHub under Apache-2.0.
This model occupies a specific category: decision models, sometimes called system one models. They sit between LLMs and traditional classifiers. They are faster and more capable at a given size than an LLM for bounded decisions, but they cannot generate arbitrary text. Strands Decider 2B competes with TypeSafe's Jev and with fine-tuned BERT-style classifiers. It excels at the rote decisions inside agentic workflows: model routing, tool selection, argument checking, triage, guardrails, evaluations, and any place where you need a fast answer from a fixed set of options with a reliability score.
If you need a chatbot, a coding assistant, or a summarizer, this is the wrong model. If you need to ask "Does this convey urgency?" or "Which team should handle this?" thousands of times per second on local hardware, it is a strong fit.
Strands Decider 2B is a dense model with 2 billion parameters. Every parameter is active for each inference. The base is Qwen3.5-2B-Base. On top of that, AWS adds a LoRA adapter and a small readout head. The readout head scores the options of a typed question. The model does not generate tokens autoregressively. It processes the state and the question in a single forward pass and returns a probability distribution over the allowed answers. That design is why latency is low and why throughput is high.
The 4,096 token context is enough for a prompt, a tool description, a short conversation history, or a chunk of text to classify. It is not enough for long documents or multi-turn reasoning. The model outputs a calibrated confidence with every answer. On short classification tasks it has never seen, answers at a confidence of 0.9 or more are right about 95% of the time. Below that, you should confirm or ask a person.
Training data includes many public classification datasets: BoolQ, AG News, DBpedia, SST-5, emotion, banking77, Amazon Massive Intent, PAWS, Civil Comments, HelpSteer2, HotpotQA, and others. The training recipe, data inventory, and evaluation scripts are in the GitHub repository.
This model answers three types of questions. Each one returns a calibrated confidence.
Concrete use cases for local agents:
It is not for coding, chat, summarization, or any task that requires generating text. It is for fast, bounded decisions.
The model is small enough to run on a laptop, a mini PC, or a single consumer GPU. Because it is a 2B dense model, the memory footprint is modest. The official checkpoint is a LoRA adapter on Qwen3.5-2B-Base, so you need the base weights plus the adapter and head. The adapter and head add very little memory.
Approximate VRAM requirements for the base weights:
For CPU-only inference, 8 GB of system RAM is enough. The model returns answers in tens of milliseconds on a modern CPU. On a GPU, expect sub-10ms per decision. On an RTX 4090, throughput can reach hundreds to thousands of decisions per second depending on batch size. On an M4 Max, use --device mps and expect similar low latency. On an RTX 3060 12GB, you can run FP16 comfortably. The best GPU for Strands Decider 2B is any modern GPU with 6 GB or more VRAM.
A note on "tokens per second": this model does not generate tokens. It scores options in one pass. If you see tokens per second benchmarks for Qwen3.5-2B, that is the base LLM, not the decider. The right metric here is decisions per second or milliseconds per decision.
Recommended quantization: Q4_K_M is the best starting point for most users who are memory constrained. It fits in 3-4 GB VRAM and keeps latency low. The tradeoff is that lower quantization can degrade the calibration of the confidence scores. If you rely on the 0.9 confidence threshold, test Q8 or FP16 first.
The quickest way to get started is the official CLI, not Ollama. Ollama does not natively support this model because it is a LoRA adapter plus a custom readout head. The official path is:
1pip install strands-decider2strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \3 --state "Help! My payouts have been failing for 3 days!" \4 --choice "Which team should handle this?=billing,sales,retail" \5 --noul "Does this convey urgency?" \6 --score "How frustrated is the writer?=calm,frustrated,depressed"
Add --device cuda, --device mps, or --device cpu to force a device. The server mode binds to HTTP if you want to call it from other services. If you want to run Strands Decider 2B on an RTX 4090, use --device cuda. On an M4 Max, use --device mps. On a CPU-only box, the default will pick CPU.
Strands Decider 2B vs Qwen3.5-2B: They share the same base. Qwen3.5-2B is a generative LLM. It can chat, write code, and summarize. Strands Decider 2B cannot do any of that. But for typed decisions with calibrated confidence, Strands Decider is faster and more reliable at the same size. Choose Qwen3.5-2B if you need text generation. Choose Strands Decider 2B if you need a fast answer from a fixed set of options.
Strands Decider 2B vs Gemma 2 2B: Gemma 2 2B is another small general-purpose LLM. It is better for open-ended generation and reasoning. Strands Decider 2B is better for classification, scoring, and routing. If your workload is "pick one of these five labels and tell me how sure you are," Strands Decider is the more efficient tool.
Strands Decider 2B vs TypeSafe's Jev: Jev popularized the decision model category. Strands Decider 2B is fully open source under Apache-2.0, with published training data, scripts, and evaluations. It briefly reached the top spot on the JevBench ranking for models of its size. If you need an open, auditable decision model that runs locally, Strands Decider 2B is a strong choice. If you need a hosted API with a commercial SLA, Jev may be the better fit.
For local AI in 2026, a 2B parameter decision model is a practical building block. It will not replace your LLM. It will handle the small, repetitive decisions that would otherwise waste LLM tokens and add latency to your agent.

Explore the Provider
Aggregate stats, leaderboard, release timeline, and benchmark coverage across every Amazon Web Services model we track.

Check which of your devices can run this model and how fast.

Compare hosted per-token prices before you commit to local hardware.

Break-even math for self-hosting, renting a GPU, or paying per token.
Live status for the hosted APIs you might use instead of running this locally.