Microsoft to announce MAI-Thinking-1 reasoning model and Windows 11 developer mode at Build
Microsoft will announce MAI-Thinking-1 at its Build conference on June 2, 2026, according to sources cited by The Verge. The model is Microsoft's first reasoning model and was not trained using distillation from other AI models. The company will also reveal MAI-Image-2.5 and MAI-Image-2.5-Flash image models, along with a new developer-optimized Windows 11 experience.
Microsoft to announce MAI-Thinking-1 reasoning model and Windows 11 developer mode at Build
Microsoft will unveil MAI-Thinking-1, its first reasoning model, at the Build conference on June 2, 2026, according to sources cited by The Verge's Tom Warren. The model was not trained using distillation — meaning it wasn't trained by learning from another AI model's outputs — and will target enterprise use cases.
Microsoft AI chief Mustafa Suleyman will present the model alongside announcements of MAI-Image-2.5 and MAI-Image-2.5-Flash, which Suleyman teased last week. Pricing, context windows, and benchmark scores have not been disclosed.
Windows 11 developer experience
Microsoft plans to announce a new Windows 11 developer-optimized experience featuring a distraction-free environment with pre-installed development tools, apps, and scripts. The company will also detail efforts to rewrite parts of Windows 11 for improved performance, building on improvements already rolling out to Windows Insiders.
The conference will emphasize local model execution on Windows, allowing developers to run AI models on-device rather than relying on cloud services. This includes support for Nvidia's RTX Spark silicon, which Nvidia announced as "the most efficient PC chip ever built" at Computex.
Copilot super app in development
Microsoft is building a Copilot "super app" that combines its various Copilot AI assistants into a single interface, according to sources. The app will include Microsoft Scout, an AI agent reportedly based on Microsoft's OpenClaw work. A leaked screenshot circulating on Friday was a mockup prepared for Build demonstrations. The app won't be available at Build, with a late summer preview expected.
Conference context
Build 2026 is taking place in a smaller, more intimate San Francisco venue as Microsoft attempts to rebuild developer trust following GitHub departures, outages, and security incidents. Satya Nadella will discuss the RTX Spark announcement with Nvidia CEO Jensen Huang during his keynote, with Qualcomm also expected to present on Windows on Arm improvements.
The keynote begins at 9:30 AM PT on June 2.
What this means
Microsoft's non-distilled reasoning model represents a different training approach from competitors like OpenAI's o1 series and DeepSeek-R1. Enterprise targeting suggests Microsoft sees reasoning models as differentiated enough to warrant dedicated business use cases rather than general consumer deployment. The emphasis on local Windows execution aligns with broader industry movement toward on-device AI, though actual performance and cost comparisons against cloud alternatives remain to be demonstrated.
Related Articles
Tencent Open-Sources Hy4 Preview: 770B-Parameter MoE Model with 1M-Token Context
Tencent's Hy Team has open-sourced Hy4 preview, a 770-billion-parameter Mixture-of-Experts model with 49 billion activated parameters and a 1-million-token context window. The model is available under Apache 2.0 alongside an FP8-quantized variant, with Tencent claiming it beats GLM 5.3 and Kimi K3 on internal engineering evaluations.
Tencent Releases Hy4 Preview: 770B-Parameter MoE Model with 1M Context for Coding Agents
Tencent has released Hy4 preview, a mixture-of-experts model with 770B total parameters and 49B active parameters, targeting coding agents and multi-step tool-use workflows. The model ships with a 1 million token context window and is priced at $0.834 per 1M input tokens and $2.501 per 1M output tokens.
Google Launches Gemini 3.5 Transcribe with 4.0% Word Error Rate Across 85 Languages
Google has released Gemini 3.5 Transcribe, a speech-to-text model that automatically detects 85 languages, removes filler words, and corrects misspoken phrases. The company claims a 4.0 percent word error rate for streaming audio and 70 percent lower latency than its predecessor, Chirp 3.
Z.ai's GLM-5.3-Flash Matches Top Models at 7.5x Lower Cost, Runs Entirely on Chinese Chips
Z.ai released GLM-5.3-Flash, a 320-billion-parameter MoE model with an 18-billion active parameter count and a one-million-token context window. It nearly matches the larger GLM-5.3 on Artificial Analysis's Intelligence Index while costing roughly 7.5 times less per task, and it reportedly runs entirely on Chinese AI chips instead of Nvidia GPUs.
Comments
Loading...