Mistral releases Devstral Medium and Small 1.1 with 61.6% SWE-Bench Verified score
Mistral AI has released two specialized coding models: Devstral Medium, achieving 61.6% on SWE-Bench Verified, and Devstral Small 1.1, scoring 53.6% and released under Apache 2.0 license. The company claims Devstral Medium surpasses Gemini 2.5 Pro and GPT-4.1 at a quarter of the price.
Devstral Medium — Quick Specs
Mistral releases Devstral Medium and Small 1.1 with 61.6% SWE-Bench Verified score
Mistral AI has released two specialized coding models developed in collaboration with All Hands AI: Devstral Medium and Devstral Small 1.1. The models are designed specifically for agentic coding tasks, with emphasis on generalization across different prompts and agentic scaffolds.
Devstral Medium: API-only proprietary model
Devstral Medium achieves 61.6% on SWE-Bench Verified, according to Mistral AI. The company claims the model surpasses Gemini 2.5 Pro and GPT-4.1 at a quarter of the price, though specific benchmark comparisons were not provided.
Pricing for Devstral Medium (devstral-medium-2507):
- Input: $0.40 per 1M tokens
- Output: $2.00 per 1M tokens
The model is available through Mistral's API and supports on-premise deployment for enterprise customers. Custom fine-tuning is available for enterprises requiring task-specific optimization.
Devstral Small 1.1: Open-source Apache 2.0 release
Devstral Small 1.1 scores 53.6% on SWE-Bench Verified. Mistral claims this sets a new state-of-the-art for open models without test-time scaling, though the model maintains the same 24B parameter architecture as its predecessor.
Pricing for Devstral Small 1.1 (devstral-small-2507):
- Input: $0.10 per 1M tokens
- Output: $0.30 per 1M tokens
Key improvements over the previous version:
- Enhanced performance on SWE-Bench Verified (previous score not disclosed)
- Better generalization to different coding environments
- Support for both Mistral function calling and XML formats
- Optimized for use with OpenHands agentic framework
Technical specifications
Both models support:
- Multiple agentic scaffolds and prompting formats
- Integration with coding environments
- Function calling capabilities
Devstral Small 1.1 is released under the Apache 2.0 license, allowing unrestricted commercial and research use. The model is available for local deployment. Devstral Medium remains proprietary but can be deployed on private infrastructure through enterprise agreements.
Context window size, training data cutoff, and detailed architecture specifications were not disclosed.
What this means
Mistral is positioning itself in the increasingly competitive coding model space with a two-tier strategy: an open-source model for local deployment and experimentation, and a proprietary API model targeting enterprise customers. The 61.6% SWE-Bench Verified score for Devstral Medium, if independently verified, would be competitive with leading coding models, though claims of cost advantage over Gemini 2.5 Pro and GPT-4.1 require context on benchmark parity. The Apache 2.0 release of the 24B parameter Small model provides the open-source community with a capable coding agent foundation without licensing restrictions.
Related Articles
Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October
Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.
Ai2 open-sources AstaBrief 8B, a Qwen3-8B report model it says runs 3.5x faster than Claude in Asta
Ai2 has open-sourced AstaBrief 8B, a model fine-tuned from Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. It is live in Asta as Fast mode, which averages 51.1 seconds per report versus 178.5 seconds for the Claude-powered Thinking mode, according to Ai2. The weights and training data are public.
Microsoft releases Decision-1, a Qwen3.5-9B-based model for classification and routing, at $0.042 per 1M input tokens
Microsoft has released Decision-1, a decision model built on Qwen3.5-9B for classification, evaluation, and routing. Microsoft claims 83.5% accuracy across 36 benchmarks and 85 ms latency. Input tokens cost $0.042 per million, and output tokens are free.
Microsoft releases FrogNano-4B, an Apache 2.0 coding agent trained with RL on 1,500 synthetic tasks
Microsoft has released FrogNano-4B-2609, a repository-level coding agent derived from Qwen3.5-4B and published under Apache 2.0 with open weights. Microsoft says it was post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. It is evaluated at roughly 131K tokens of context.
Comments
Loading...