Mistral Releases Medium 3.5: 128B Model with Cloud Coding Agents and 77.6% SWE-Bench Verified
Mistral AI released Medium 3.5, a 128B dense model with a 256k context window that scores 77.6% on SWE-Bench Verified. The model powers new remote coding agents in Mistral Vibe that run asynchronously in the cloud, plus a new Work mode in Le Chat for multi-step agentic tasks.
Mistral Medium 3.5 — Quick Specs
Mistral Releases Medium 3.5: 128B Model with Cloud Coding Agents and 77.6% SWE-Bench Verified
Mistral AI released Mistral Medium 3.5, a 128B dense model with a 256k context window, built for long-running coding and productivity tasks. The model scores 77.6% on SWE-Bench Verified and 91.4 on τ³-Telecom, according to Mistral.
Mistral Medium 3.5 is now available in public preview under a modified MIT open-weights license. The company claims the model can run self-hosted on as few as four GPUs and combines instruction-following, reasoning, and coding in a single set of weights.
Performance and Technical Details
According to Mistral, the model outperforms Devstral 2 and Qwen 3.5 397B A17B on SWE-Bench Verified with its 77.6% score. The model includes a vision encoder trained from scratch to handle variable image sizes and aspect ratios.
Reasoning effort is configurable per request, allowing the same model to handle quick chat replies or complex agentic workflows. Mistral built the model specifically for long-horizon tasks with reliable tool calling and structured output.
Pricing and Availability
Mistral Medium 3.5 is priced at $1.50 per million input tokens and $7.50 per million output tokens via API. Open weights are available on Hugging Face. The model is also available through NVIDIA's build.nvidia.com platform and as NVIDIA NIM containerized inference microservices.
Remote Coding Agents in Vibe
Mistral launched remote coding agents in Vibe CLI that run asynchronously in the cloud. Users can start coding sessions from either the CLI or Le Chat web interface, with sessions running in isolated sandboxes while users step away.
Local CLI sessions can be "teleported" to the cloud, preserving session history, task state, and approvals. The agents integrate with GitHub for pull requests, Linear and Jira for issues, Sentry for incidents, and Slack or Teams for notifications.
According to Mistral, the system is designed for high-volume coding work like module refactors, test generation, dependency upgrades, and bug fixes. Multiple coding sessions can run in parallel.
Work Mode in Le Chat
Mistral introduced Work mode in Le Chat (Preview), powered by Medium 3.5, for multi-step agentic tasks beyond coding. The mode enables cross-tool workflows, research and synthesis, inbox triage, and issue creation in project management tools.
In Work mode, connectors are enabled by default, allowing the agent to access documents, mailboxes, calendars, and other systems. Every tool call and reasoning step is visible to users, with explicit approval required for sensitive actions like sending messages or modifying data.
What This Means
Mistral's combination of a strong open-weights coding model with cloud-based agent infrastructure directly addresses the friction in current AI coding workflows, where developers must babysit agent runs. The 77.6% SWE-Bench Verified score positions Medium 3.5 competitively against larger models, while the claimed four-GPU deployment requirement could enable wider self-hosting.
The $1.50/$7.50 pricing undercuts similar-capability models from competitors, though real-world performance on complex codebases will determine adoption. The integration of coding agents into Le Chat and the new Work mode signals Mistral's push beyond chat interfaces into persistent, multi-step autonomous workflows.
Related Articles
Microsoft releases FrogNano-4B, an Apache 2.0 coding agent trained with RL on 1,500 synthetic tasks
Microsoft has released FrogNano-4B-2609, a repository-level coding agent derived from Qwen3.5-4B and published under Apache 2.0 with open weights. Microsoft says it was post-trained only with reinforcement learning on about 1,500 synthetic software-engineering tasks, with no stronger-model trajectories. It is evaluated at roughly 131K tokens of context.
StepFun releases Step 5 Preview: 600B MoE with 1M context at $1/$2.70 per 1M tokens
StepFun has listed Step 5 Preview, a sparse Mixture-of-Experts model with 600B total and 27B active parameters and a 1.0M-token context window. It is priced at $1 input and $2.70 output per 1M tokens on OpenRouter. StepFun positions it as its flagship model for agentic work.
OpenAI launches GPT-6 in ChatGPT with 'Intelligent UI' and interactive answers; Sol for paid users, Luna for free
OpenAI is rolling out GPT-6 to all ChatGPT tiers, with paying users on GPT-6 Sol and free users on GPT-6 Luna. The release adds 'Intelligent UI,' which renders answers as interactive charts, buttons, forms and mini apps, and lets the model respond while still thinking. OpenAI claims this cuts wait times by 44 percent.
Liquid AI releases open d1-3B decision model: 16 ms on Jetson AGX Thor, 48.57 on Decision Index
Liquid AI released two open-weight decision models, d1-3B (text and image) and the experimental d1-omni-600M (text with image or audio). Unlike generative models, they answer in a single forward pass, and Liquid AI claims d1-3B scores 48.57 on its Decision Index 0.2.1, ahead of all 4B and 9B models it tested.
Comments
Loading...