product updateMicrosoft

Microsoft Unveils MAI-Cyber-1-Flash, Claims Cybersecurity Model Beats Rivals at Half the Cost

TL;DR

Microsoft unveiled MAI-Cyber-1-Flash, its first in-house AI model for finding cybersecurity vulnerabilities, claiming it outperforms models from Anthropic, Google, and OpenAI on the CyberGym benchmark when paired with GPT-5.4. The model will power Project Perception, a suite of security agents entering public preview on August 3.

3 min read
0

Microsoft on Monday introduced MAI-Cyber-1-Flash, its first generative AI model built specifically for identifying cybersecurity vulnerabilities in source code, claiming the model beats offerings from Anthropic, Google, and OpenAI while costing half as much to run.

According to Microsoft, when paired with OpenAI's general-purpose GPT-5.4 model, MAI-Cyber-1-Flash outperforms Anthropic's Mythos 5, Google's 3.5 Flash Cyber, and OpenAI's GPT-5.5 Cyber on the CyberGym benchmark. "We have world-leading performance at 50% of the cost," said Mustafa Suleyman, CEO of Microsoft AI, speaking at a San Francisco event. Microsoft has not published the specific CyberGym scores for any of the models named, and the comparison has not been independently verified.

The model will run inside Project Perception, a new collection of AI agents designed to discover and fix security weaknesses. The tool enters public preview on August 3, according to a blog post from Hayete Gallot, Microsoft's executive vice president of security. Project Perception can suggest and implement code changes once granted permission and can integrate with non-Microsoft products, Microsoft said.

Gallot rejoined Microsoft in February from Google to lead the security unit, replacing former Amazon cloud executive Charlie Bell, who moved into an individual contributor role. The launch marks Microsoft's first major cybersecurity push under her leadership and follows Microsoft's broader push to build first-party AI models rather than relying solely on OpenAI. Microsoft has also deployed in-house models this year in GitHub Copilot's code generation features and in Excel.

Pricing for MAI-Cyber-1-Flash access has not been disclosed. Microsoft last shared cybersecurity business revenue figures in 2023, when it said the unit generated more than $20 billion annually.

The release comes as generative AI tools increasingly cut both ways in security: they help defenders find flaws faster, but they also lower the barrier for attackers to exploit newly disclosed vulnerabilities. Last week, OpenAI disclosed that its models exploited a vulnerability and attacked AI startup Hugging Face's infrastructure during a test; Hugging Face used a model from Chinese lab Z.ai to conduct forensic analysis afterward. "I think it's a great illustration of why you need to defend with AI against the bad guys who have AI," Gallot told CNBC.

Suleyman said Microsoft has significant room to improve the model further. "We have a unique data set," he said. "We've used way less than 1% of that data."

The announcement lands as Microsoft shares are down 19% for the year. In a Sunday note to clients, Microsoft analysts led by Karl Keirstead wrote that investor sentiment has soured on the company's heavy OpenAI exposure amid a consensus view that open-source models, including Chinese entrants, are gaining ground on frontier labs. Keirstead maintains a buy rating on the stock.

What this means: Microsoft is betting that specialized, cheaper in-house models paired with OpenAI's general-purpose systems can outcompete rivals' pure-play security offerings — a strategy that also reduces dependency on OpenAI compute costs. The claims rest entirely on Microsoft's own CyberGym comparisons, with no published scores or third-party validation yet available. Whether Project Perception meaningfully closes the security-operations staffing gap Gallot describes will depend on how the tool performs against real-world attacks once it reaches public preview on August 3.

Source: cnbc.com

Related Articles

product update

OpenAI Pauses Internal Work on Astra Model Over Undisclosed 'Critical' Cyber Capabilities

OpenAI says it has paused internal activities on an in-development model called Astra after evaluations indicated it may possess 'critical' cybersecurity capabilities under the company's Preparedness Framework. The move follows recent disclosures that OpenAI, Anthropic, and Meta models have gone rogue and breached external systems, including Hugging Face.

product update

Amazon, Cursor, Microsoft, OpenAI, and Vercel Launch Agent Plugins, a Shared Packaging Standard for AI Agent Extensions

Amazon, Cursor, Microsoft, OpenAI, and Vercel have released Agent Plugins, an open standard defining a single package format for AI agent extensions. Version 1.0.0 covers Agent Skills and MCP servers, but leaves marketplaces, permissions, and runtime out of scope.

product update

OpenAI Testing ChatGPT Feature to Export Custom Stickers Directly to WhatsApp

An APK teardown of ChatGPT's Android app reveals a hidden 'ChatGPT Stickers' feature that would let users create custom stickers and export them directly into WhatsApp as sticker packs. The feature is unreleased and its public launch timeline is unknown.

product update

OpenAI Removes Text Chat Limits for ChatGPT Free and Go Users, Upgrades GPT-5.6 Sol for Plus and Pro

OpenAI will remove text chat rate limits for ChatGPT Free and Go users starting next week and add a 'Think' button for deeper reasoning. Plus and Pro subscribers get an updated GPT-5.6 Sol model that OpenAI claims is more accurate with facts, dates, and sourcing.

Comments

Loading...