Microsoft releases Decision-1, a Qwen3.5-9B-based model for classification and routing, at $0.042 per 1M input tokens
Microsoft has released Decision-1, a decision model built on Qwen3.5-9B for classification, evaluation, and routing. Microsoft claims 83.5% accuracy across 36 benchmarks and 85 ms latency. Input tokens cost $0.042 per million, and output tokens are free.
Microsoft has released Decision-1, a small model built on Qwen3.5-9B for fast, structured decisions. It is priced at $0.042 per 1M input tokens, with output tokens free, and is available through Microsoft Foundry and OpenRouter.
What Decision-1 does
Decision-1 targets three task types: classification, evaluation, and routing decisions. According to Microsoft, the model has "the potential" to "guide and control agents through complex environments." That agent-control use is framed as potential, not a demonstrated capability.
The model joins a category that has grown quickly since mid-September. Decision models from OpenAI and Cloudflare are already on the market, along with the "original" from startup Jev. All are built for fast, structured outputs rather than open-ended generation.
Specifications
| Item | Detail |
|---|---|
| Base model | Qwen3.5-9B (about 9B parameters) |
| Input price | $0.042 per 1M tokens |
| Output price | Free ($0) |
| Availability | Microsoft Foundry, OpenRouter |
| Context window | Not yet disclosed |
| Training cutoff | Not yet disclosed |
Benchmark claims
Microsoft says Decision-1 is the most accurate model tested across 36 benchmarks covering nearly 150,000 questions. In Microsoft's own comparison, it scores 83.5% accuracy with 85 ms latency, ahead of Jev 1.13.0.
On speed, Microsoft claims Decision-1 is 2.5 times faster than the runner-up, H2O-Lightning-4B.
These figures come from Microsoft's own testing and have not been independently verified. The comparison also leaves out Cloudflare's open-source Clef models, which are likewise based on Qwen. Without that head-to-head, the "most accurate" claim covers only the models Microsoft chose to include.
Pricing
At $0.042 per million input tokens with free output, cost is driven almost entirely by input volume. Decision tasks typically produce short outputs such as a label, a score, or a route. Free output tokens therefore matter less than the low input rate, but they simplify cost estimates for high-volume pipelines.
Market context
Jev started the decision model trend in mid-September. According to The Decoder, the approach was quickly adapted and surpassed using open small language models. Decision-1 fits that pattern: a Qwen base model specialized for a narrow task and distributed through standard channels.
What this means
Decision models are becoming a commodity layer. Within weeks of Jev's launch, Microsoft, OpenAI, and Cloudflare have all shipped competing offerings. Because Decision-1 and Cloudflare's Clef both build on Qwen, the differentiators are fine-tuning data, evaluation design, and distribution, not base architecture.
For engineers building agent systems, the practical question is whether a 9B specialist can replace a call to a frontier model for routing and triage. Microsoft's 85 ms latency claim suggests it can in latency-sensitive paths, but the missing context window figure and the absence of Clef from the comparison leave gaps. Teams should run Decision-1 against their own labeled routing data before switching.
The benchmark methodology also matters. A 36-benchmark aggregate can hide weak performance on specific task types, so per-benchmark results are needed to judge fit for a given workload.
Related Articles
Microsoft releases FrogNano-4B, an Apache 2.0 coding agent trained with RL on 1,500 synthetic tasks
Microsoft has released FrogNano-4B-2609, a repository-level coding agent derived from Qwen3.5-4B and published under Apache 2.0 with open weights. Microsoft says it was post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. It is evaluated at roughly 131K tokens of context.
StepFun releases Step 5 Preview: 600B MoE with 1M context at $1/$2.70 per 1M tokens
StepFun has listed Step 5 Preview, a sparse Mixture-of-Experts model with 600B total and 27B active parameters and a 1.0M-token context window. It is priced at $1 input and $2.70 output per 1M tokens on OpenRouter. StepFun positions it as its flagship model for agentic work.
Microsoft's Copilot gets access to local Windows files and OS-level actions under 'Hybrid Intelligence'
Microsoft announced an upgrade to Copilot at its Windows and Surface event that gives the assistant access to local files and the ability to take actions across Windows. The company calls the underlying approach "Hybrid Intelligence," which combines local and cloud AI models. Pricing, model details, and availability were not disclosed in the available reporting.
Microsoft will turn Windows Search into a Copilot-connected command interface this fall
Microsoft announced a redesigned Windows Search menu that accepts short typed commands to change system settings and can converse with the new Copilot app without launching it. The company says the update arrives this fall. Model, pricing and availability details were not disclosed.
Comments
Loading...