AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Options in About 115 ms
AWS Strands Labs releases Strands Decider 2B, an open source decision model. It does not generate text. It reads a state and typed questions, then returns a choice, a yes/no probability, or a score with a calibrated confidence. The model has 1.9 billion parameters and runs locally on a CPU, a consumer GPU, or an Apple silicon Mac.
Is it deployable? Yes, for local and self-hosted use. Weights are on Hugging Face under Apache-2.0, and pip install strands-decider gives a CLI and an HTTP server. The bundled server binds to 127.0.0.1 with no authentication, so production needs your own auth layer. No hosted inference provider serves it yet.
What a decision model does
Decision models, also called System One models, became a category after TypeSafe AI launched Jev last month. An LLM can produce arbitrary output. A decision model only picks between options or rates on a scale.
Strands Decider supports 3 question types:
- choice: pick 1 of N options.
- noul: a yes/no probability between 0 and 1.
- score: a level on an ordered rubric.
Every answer comes from the allowed options and carries a confidence. The team states the model is worse than reasoning models on complex problems. It is unsuited for coding, chat, or summarization.
Architecture: an LLM with its mouth removed
The team starts from Qwen3.5-2B-Base and discards the language-modelling head. A small pointer head of about 1 million parameters replaces it. That head compares the hidden state at the <answer> position against the hidden state at each option’s last token. One forward pass yields the result, with no decoding loop.
The torso uses a rank-16 LoRA, and the head runs in fp32. Label sets come from the request, so nothing caps the option count. The released checkpoint is v19.
Asking several questions about one text is cheap. The state is read once, and each extra question adds only its own tokens.
Benchmarks and latency
The team measures accuracy and calibration on the public set of JevBench, a third-party benchmark for Jev-class models. Published v19 figures:
- JevBench v1 public accuracy: 0.723 (167 of 231 tasks).
- Brier score 0.342, expected calibration error 0.052.
- Tier accuracy: easy 1.000, standard 0.875, hard 0.505.
- Latency on an RTX 3090: 115 ms median, 299 ms p95.
- Latency on an M3 Pro: 153 ms warm median under 300 tokens.
On the v1.4.2 board of September 25, v19 ranked 3rd of 33 in the 2B class. Excluding 3 models just over 2B, it ranked 1st of 30. The repo also flags a caveat. Mapika’s newer decider-2b v11 scores 175 of 231 on the Strands harness, 8 tasks ahead. Strands Decider was not on the newer v1.5.4 composite board at the time of writing.
Calibration is the practical win. On unseen short classification tasks, answers at 0.9 confidence or higher were right about 95% of the time. The team advises confirming or escalating below that threshold.
How it compares
| Feature | Strands Decider 2B (v19) | Jev 1.13.0 | decider-2b | Decision 2B |
|---|---|---|---|---|
| Developer | Strands Agents (AWS) | TypeSafe AI | Mapika | FlyMy.AI |
| Access | Open weights, Apache-2.0 | Closed hosted API | Open code or weights | Open code or weights |
| Base model | Qwen3.5-2B-Base + LoRA + pointer head | Undisclosed | Qwen3.5-2B-Base + trained readout | MiniCPM5-2B + LoRA + pointer head |
| Size | 1.9B | Undisclosed | 1.9B | 2.5B dense |
| Self-hosting | Yes | No | Yes | Yes |
| Full training recipe and data list published | Yes | No | Not verified | Not verified |
| JevBench public accuracy (v1.4.2 board) | 0.723 | Not in source table | 0.710 | 0.753 |
| Reported latency | 115 ms median (RTX 3090) | 70 to 500 ms (vendor) | Not compared | Not compared |
Sources: Strands JevBench comparison, Benchmark Heaven JevBench, Jev product page. Latencies come from different hardware and harnesses, so they are not directly comparable.
Use cases and a guardrail example
The team reports early success in model routing, tool selection, argument checking, triage, guardrails, evals, and hybrid agents. In a hybrid agent, an LLM makes the hard calls and the decider handles rote ones.
The repo example gates a weather tool call inside a Strands agent. A before_tool_call intervention asks 2 yes/no questions. Are the arguments grounded in what the user said? Is calling now premature? If the agent guessed a city, it asks the user instead.
From the CLI, routing “Help! My payouts have been failing for 3 days!” across billing, sales and retail returns billing with confidence 0.768.
Key Takeaways
- Strands Decider 2B returns choices, yes/no and scores, never text.
- It swaps Qwen3.5-2B’s LM head for a ~1M-parameter pointer head.
- v19 scores 0.723 on JevBench public with ECE 0.052.
- Median latency is 115 ms on an RTX 3090.
- Weights, code, data list and recipe ship under Apache-2.0.
Check out the technical details, GitHub repo and model weights. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us