Strands has released a two-billion-parameter open-source model designed to choose among supplied options rather than generate unrestricted text. Called Strands Decider 2B, the model targets fast, repeatable decisions inside software agents, including tool selection, routing, evaluations, guardrails and policy classification.

The architecture starts with the torso of Qwen3.5-2B but removes its language-model head. In its place, Strands uses a pointer head of slightly more than one million parameters to score answer options against a designated answer position. The underlying torso is adapted with a rank-16 LoRA layer. Model weights, training data and build scripts have been published through GitHub and Hugging Face, according to the project.

That structure creates a sharply defined tradeoff. A developer can ask the model to identify a language, classify sentiment or select a tool from a list, and the system will return one of the allowed choices with probability-like reliability scores. It cannot draft arbitrary text, summarize documents, write code or serve as a general chatbot. Complex reasoning problems are also outside its intended role.

Strands says the limited output space improves latency and makes it efficient to ask several questions about the same input. The company reports median local decision time of about 115 milliseconds on an Nvidia RTX 3090 for its tested workloads. Small tasks took about 153 milliseconds on an M3 MacBook, with latency increasing roughly in proportion to task size. These are developer-reported measurements rather than independent benchmarks.

On the public portion of JevBench, Strands reports that the model ranked third among 33 entries in the 2B class when accuracy and calibration were considered, and first among 30 after excluding models slightly above two billion parameters. It also completed all tasks categorized as easy correctly. The released model is version 19, following experiments with an earlier slot-head design that the team says performed worse.

The project positions decision models as companions to, not replacements for, large language models. A hybrid agent could send routine classifications to a local decider while reserving a larger generative model for ambiguous requests. That could reduce inference cost and response time, provided developers carefully define the available choices and recognize when a task needs broader reasoning.

Strands Decider can run on a local CPU or GPU and is accessible through a command-line tool. Examples in the repository show it working inside a Strands agent alongside an external language model. Its open training materials may be as important as the initial scores, giving developers a reproducible base for testing where constrained decision systems fit into agent architectures.