AI NewsModels & agentsAnnouncement
AWS released Strands Decider 2B, an open-source decision model that returns 115 ms picks on a desktop GPU
AWS Strands Labs released Strands Decider 2B, a 2-billion-parameter open-source decision model that returns typed picks in about 115 milliseconds on a desktop RTX 3090 and in about 153 milliseconds on an M3 MacBook. Weights are on Hugging Face and the code is on GitHub.

Why it mattersA team routing an agent, choosing a tool or guarding a model with a classifier now has a free, local option that AWS says ranks third of thirty-three models of its size on an independent benchmark.
Picking the next step for an agent, deciding which tool to call, choosing which memory to load, deciding whether a message is safe to send: these are classifier jobs, and until this month the common solution was a prompt into a language model that cost tokens and seconds for each decision. AWS Strands Labs shipped its own small, open-source decision model today, Strands Decider 2B, and reports that it returns a pick in about 115 milliseconds on a desktop GPU.
AWS's blog post places the release inside a wider set of launches: TypeSafe AI's Jev, OpenAI's Decisions API, Cloudflare's Clef and Ollama's new decision endpoint all landed in the past two weeks. The pitch for Strands Decider is open weights and a small enough footprint to run locally.
What the model is and how it works
The model has 2 billion parameters and is built from the "torso" of Qwen 3.5-2B. AWS replaced the normal language-modelling head, which generates one word at a time, with what it calls a pointer head that scores every answer in parallel. The decision is one forward pass through the backbone, with no autoregressive generation. On top of that, AWS added a rank-16 LoRA adapter, so the fine-tune touches only a small set of weights.
The weights are on Hugging Face under the StrandsAgents account, the code is on GitHub at strands-labs/strands-decider, and AWS lists pip install strands-decider as the install path. The post names Marc Brooker, Mike Chambers and Fabio Nonato de Paula as the three authors.
The numbers AWS reports
AWS's own measurement on the public JevBench ranking: Strands Decider 2B is "3rd of 33 in the 2B class, and 1st of 30 excluding the just-over-2B models". Latency, measured on hardware AWS names, is about 115 milliseconds on an Nvidia RTX 3090 and about 153 milliseconds on an M3 MacBook. AWS says latency rises roughly in a line with task size.
Each one of those is AWS's own number. The benchmark it cites, JevBench, is a third-party ranking published by Benchmark Heaven, so the position on the ladder is independent of the figures AWS measured for its own latency.
The decision for a team
Strands Decider is the fifth decision model to land on the market in two weeks, which is the shape of the pattern: Jev, Clef, Decisions API, Ollama's endpoint and now this. The practical question for a team is not which one is best on a vendor's own chart, which is a hard thing to compare across posts, but which one is close enough to the task that one week of eval work on the team's own data proves the swap. A team already paying for a frontier model to classify every ticket, score every retrieval or route every tool call will see the cost line change the day the swap goes live. The cheapest proof is to replay a sample of production traffic against the new model and read the confusion matrix, before deciding where it is safe to put in a live production path. For a team that cares about the licence or wants the model on its own hardware, the open-weights option is now AWS, Cloudflare, Ollama and a handful of Hugging Face projects.
Source
- Primary source: Strands Agents: Introducing Strands Decider, 1 October 2026
- Press coverage: The New Stack: AWS launches a local answer to TypeSafe's Jev decision model, 1 October 2026
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

