Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base
Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.
Amazon Web Services has released Strands Decider 2B, a fully open-source, 2-billion-parameter decision model that selects among pre-defined options and returns a measure of its confidence in each choice. The model is available now and is small enough to run locally, according to TechCrunch's reporting.
It comes from Strands Labs, the AWS organization building tools and protocols for deploying AI agents. It was released the same week OpenAI announced a similar offering. Details of OpenAI's product were not included in the source report.
What Strands Decider does
Decision models do not generate free-form text. They classify a situation into one of a fixed set of answers and report how sure they are. Strands Decider is built on the "torso" of an LLM, in this case a Qwen3.5-2B base (the source spells it "Qen3.5-2B"), with the text-generation behavior replaced by calibrated choices.
The target use is the routing step in agentic workflows. Amazon distinguished engineer Marc Brooker, who started the project, described the use case as a decider for a workflow step: "what is the next thing for me to do here, based on where I am?" He told TechCrunch the approach gives customers a step that is "more reliable, thanks to the confidence scores, thanks to the closed domain of answers, [and is] lower latency, potentially lower cost."
Origins and claims
Brooker began the project after seeing Jev, the decision model from TypeSafe, and tried to build his own version. According to the report, the prototype briefly reached the top spot on the Jevbench ranking for models of its size. Amazon engineers then cleaned it up and released it through Strands Labs. The Jevbench claim comes from the report, and no scores were disclosed.
Brooker said the need emerged in conversations with AWS customers whose agent workflows did not always require the capability or cost of a full LLM.
Specs
- Parameters: 2B
- Base model: Qwen3.5-2B, per the report
- License and weights: described as fully open-source; specific license terms not stated in the source
- Context window: not disclosed
- Pricing: not applicable for local use; no hosted pricing disclosed
- Benchmark scores: not disclosed beyond the Jevbench ranking claim
- Training cutoff: not disclosed
Market context
TypeSafe named Jev after economist William Stanley Jevons, whose theory holds that falling costs for a resource can increase demand for it. According to the report, dozens of similar models have appeared from researchers since TypeSafe introduced the idea.
Brooker said the central engineering problem is improving accuracy and calibration on decision tasks "without degrading its performance on understanding different languages, on having the kind of knowledge it has." He does not expect frontier labs to necessarily dominate the category, since building a useful small model can cost hundreds or thousands of dollars.
TypeSafe CEO Diogo Almeida dismissed the new entrants. He said the "current batch seems more like ML people wanting to implement a cool architecture than a team deeply dedicated to making intelligence useful," and that he sees no real competition yet.
What this means
Amazon's release suggests decision models are becoming a standard component of agent stacks rather than a one-company product. Routing and branching steps in agent workflows are frequent, latency-sensitive, and have closed answer sets, so a 2B model with calibrated confidence can be cheaper and more predictable than calling a frontier LLM for each step.
The open questions are quality and trust. No public benchmark numbers accompany the release, and a briefly held Jevbench lead for a given size class is a weak signal. Calibration is the property that matters most here, because a confidence score is only useful if it is accurate enough to trigger escalation to a larger model. Builders should test it on their own workflows before relying on it.
The low cost of building these models also means differentiation will likely come from training data and calibration quality, not scale.
Related Articles
AWS Publishes Reference Architecture for Contract Intelligence Using Bedrock AgentCore and Dual Claude Models
AWS published a reference architecture showing how to combine structured data extraction with Bedrock AgentCore, dual Claude models, and Amazon Quick to answer portfolio-wide questions that standard RAG systems get wrong. The design uses Claude Sonnet 4.6 for extraction and Claude Haiku 4.5 for independent verification, with Amazon Textract as a deterministic tiebreaker.
AWS details ambient agent pattern on Bedrock AgentCore: S3 events trigger jobs, one ask_human tool pauses for approval
AWS published a reference implementation for ambient agents on Amazon Bedrock AgentCore. S3 uploads or scheduled events create jobs that an agent runs, pausing for human input through a single ask_human tool. Each agent turn is capped at the 15-minute Lambda timeout.
uniopen lifts Amazon Nova 2 Lite moderation F1 from 0.585 to 0.855 using LoRA fine-tuning on SageMaker AI
Taiwan retail platform uniopen adapted Amazon Nova 2 Lite to its two-axis moderation policy using LoRA supervised fine-tuning in Amazon SageMaker AI, plus a prompt-format change. According to AWS, Per Behavior Macro F1 rose from 0.5852 to 0.8550 and Subject Type Macro F1 from 0.4162 to 0.8491, both above production targets.
OpenAI's GPT-6.1 Sol Launches on Amazon Bedrock, Claims Near-Astra Reasoning at Fraction of Cost
OpenAI's GPT-6.1 Sol is now generally available on Amazon Bedrock, targeting agentic coding, computer use, and document-heavy business workflows. OpenAI claims the model matches GPT-6 Astra on the DeepSWE v1.1 coding benchmark at roughly one-fifth the cost per task.
Comments
Loading...