Strands releases Decider 2B, an open-source 2B decision model

Strands releases Decider 2B, an open-source 2B decision model

The Strands team has released Strands Decider 2B (strands-decider-2b), a small decision model for fast experimentation and local development. It is the team's first contribution to what it calls a new class of decision models, or system one models, a category the post says has drawn attention since TypeSafe AI launched Jev earlier this month. The code is on GitHub and the weights are on Hugging Face, along with all the training data and scripts used to build the model. The announcement follows strands-labs, which the team announced earlier this year.

Unlike an LLM, a decision model does not generate arbitrary text. It picks between a set of options and can assign simple numerical scores. The post's examples: is the string 'turn on the lights' about the coffee machine, yes or no; what language is 'sihamba ngokushesha' in, English, Zulu or Dutch; is 'this is the best doc I've ever read' positive, scored between 0 and 1. The authors say that in exchange for this reduced flexibility, decision models are faster and more capable at a given size, always produce an answer from the offered options, and can run with very low latency. They also give each decision a reliability score (how sure the model is that a yes/no is correct), which the authors say is not available through frontier LLM inference APIs, and they make it efficient to ask several questions about the same prompt.

The flip side, per the authors: generating all outputs in a single parallel pass makes these models significantly worse than reasoning models at complex problems, and with no ability to generate text they are unsuited to coding, chatbots, document summarization and other common LLM tasks.

Architecture: the team takes a pre-trained LLM torso (Qwen3.5-2B) and removes the LM head, so it can no longer generate text. In its place is a pointer head that scores the hidden state at each option position against the hidden state at the position. The head is small, just over a million parameters in total. The torso is fine-tuned with a rank-16 LoRA adapter. This is the second major iteration of the architecture; the first used a slot head that performed significantly worse. The released model is v19, and the repo documents what changed in each version.

Performance: the authors track accuracy, calibration and latency. Accuracy is measured on JevBench's public set, and calibration by Brier score on the same set. On these, strands-decider-2b ranks 3rd of 33 in the 2B class, and 1st of 30 when the just-over-2B models are excluded. It answers 100% of the easy JevBench tasks correctly. For latency, the post gives a median of around 115ms for local decisions on widely available hardware. Time grows approximately linearly with task size. The latency graph was measured on a local Nvidia RTX 3090 against v18 of the model, and the post says an M3 MacBook is not much worse, with a median around 153ms for small tasks. The authors say they have ideas to improve both, especially to lower the latency floor.

Why 2B: the team wants people to experiment, including training the model, on hardware they already have, and it sees two billion parameters as a sweet spot: small enough for experimentation, large enough for meaningful work. The authors say easy JevBench problems map well to the easier problems people give agents.

Uses: the authors report early success with model routing, tool selection, evaluations, guardrails, memory, context management and policy classification. They also describe hybrid agents, where an LLM makes the hardest decisions and a decider makes the easier rote ones to cut cost and latency, experiments combining deciders with fixed workflow languages, and people using the models to play games, automate tasks and navigate mazes.

Getting started: a strands-decider CLI lets you give the model some state and a question and get back probability scores. The repo's examples/strands/ folder has an agent example. The agent runs locally, connects to a locally running decider, and uses the default LLM from Amazon Bedrock. It has a demo get_weather tool and an eager system prompt, so when a user asks 'What's the weather?' without a place, it guesses a city. Before the tool call runs, the decider answers two yes/no questions: are the argument values grounded in anything the user said, and is it too early to call the tool. A few lines of Python turn the answers into a decision, and the agent asks which city the user meant. This uses Strands' intervention system: an InterventionHandler with a before_tool_call method, passed to Agent(interventions=[...]), returns a typed action: Proceed, Deny, Confirm (stop and ask a person) or Guide (hand the turn back to the model with feedback). The authors stress this is an illustration, not a recommendation; the questions, threshold and policy were picked by hand. The team is also working on libraries for decision model integration, with updates expected in the repo soon.

Key facts

  • Strands Decider 2B is a 2-billion-parameter open-source decision model that picks among given options and returns numeric scores; it cannot generate text.
  • It is built from a Qwen3.5-2B torso with the LM head replaced by a pointer head of just over a million parameters, fine-tuned with a rank-16 LoRA adapter; the released model is v19.
  • The authors rank it 3rd of 33 in the 2B class on JevBench accuracy and calibration (Brier score), and 1st of 30 excluding the just-over-2B models.
  • The post gives a median local decision latency of around 115ms on widely available hardware, and around 153ms for small tasks on an M3 MacBook.
  • Code is on GitHub and weights on Hugging Face, with all training data and scripts; an example agent uses it to check a tool call before it runs.

Why it matters

Decision models trade flexibility for speed and a reliability score on each answer. The authors argue that makes them a good fit for the many small yes/no and pick-one choices inside agent workflows, where an LLM call is slow and costly. Strands Decider 2B is the team's first entry in this class, which the post says has drawn attention since TypeSafe AI launched Jev earlier this month. Releasing the training data and scripts, plus a repo that documents each version up to v19, gives others a base to build on.

Who it affects

Developers building agents, especially with the Strands Harness SDK, who want cheap checks on tool calls, routing or guardrails. Researchers and hobbyists who want a small model they can run and even train on hardware they already have, such as a local CPU, GPU or an M3 MacBook. It is not for people who need a model that writes code, chats or summarizes documents; the authors say it is unsuited to those tasks.

How to use it

Get the code from GitHub and the latest weights snapshots from Hugging Face. The easiest start is the strands-decider CLI: pass it some state and a question, and it returns scores for each option, with the highest probability marking its answer. The repo's examples/strands/ folder shows the model inside a Strands agent via an InterventionHandler with a before_tool_call method, returning Proceed, Deny, Confirm or Guide. The license is not stated in the source. The team says integration libraries are coming and to watch the repo.

How solid is it

The claims come from the Strands team's own announcement. The ranking of 3rd of 33 in the 2B class (1st of 30 excluding the just-over-2B models) covers accuracy and calibration on JevBench's public set, not latency. The post does not name the models above Strands Decider 2B or give accuracy or Brier score values. The 115ms median is for 'widely available hardware', while the latency graph was measured on an RTX 3090 against v18, not the released v19. The post also says answers come in 'tens of milliseconds' and does not reconcile that with the 115ms median. No numeric latency comparison against other decision models or LLM calls is given. The accuracy and calibration claim is only that it is competitive with the other models the authors know of in this class.

Risks and caveats

The authors say the single parallel pass makes the model significantly worse than reasoning models at complex problems, and it cannot generate text. Its 100% score applies to the easy JevBench tasks only. The agent example is an illustration, not a recommendation, and its questions, threshold and policy were chosen by hand, so teams would need to design and test their own. The authors' report of early success in routing, guardrails, memory and policy classification is anecdotal and carries no figures. Latency rises roughly linearly with task size, and the authors say they still want to lower the latency floor.

“The point is that a decision this cheap can sit in a path where an LLM call never could.”

— Strands team, Strands Decider 2B announcement