Microsoft Unveils MAI-Cyber-1-Flash, Claims Cybersecurity Model Beats Rivals at Half the Cost
Microsoft unveiled MAI-Cyber-1-Flash, its first in-house AI model for finding cybersecurity vulnerabilities, claiming it outperforms models from Anthropic, Google, and OpenAI on the CyberGym benchmark when paired with GPT-5.4. The model will power Project Perception, a suite of security agents entering public preview on August 3.
Microsoft on Monday introduced MAI-Cyber-1-Flash, its first generative AI model built specifically for identifying cybersecurity vulnerabilities in source code, claiming the model beats offerings from Anthropic, Google, and OpenAI while costing half as much to run.
According to Microsoft, when paired with OpenAI's general-purpose GPT-5.4 model, MAI-Cyber-1-Flash outperforms Anthropic's Mythos 5, Google's 3.5 Flash Cyber, and OpenAI's GPT-5.5 Cyber on the CyberGym benchmark. "We have world-leading performance at 50% of the cost," said Mustafa Suleyman, CEO of Microsoft AI, speaking at a San Francisco event. Microsoft has not published the specific CyberGym scores for any of the models named, and the comparison has not been independently verified.
The model will run inside Project Perception, a new collection of AI agents designed to discover and fix security weaknesses. The tool enters public preview on August 3, according to a blog post from Hayete Gallot, Microsoft's executive vice president of security. Project Perception can suggest and implement code changes once granted permission and can integrate with non-Microsoft products, Microsoft said.
Gallot rejoined Microsoft in February from Google to lead the security unit, replacing former Amazon cloud executive Charlie Bell, who moved into an individual contributor role. The launch marks Microsoft's first major cybersecurity push under her leadership and follows Microsoft's broader push to build first-party AI models rather than relying solely on OpenAI. Microsoft has also deployed in-house models this year in GitHub Copilot's code generation features and in Excel.
Pricing for MAI-Cyber-1-Flash access has not been disclosed. Microsoft last shared cybersecurity business revenue figures in 2023, when it said the unit generated more than $20 billion annually.
The release comes as generative AI tools increasingly cut both ways in security: they help defenders find flaws faster, but they also lower the barrier for attackers to exploit newly disclosed vulnerabilities. Last week, OpenAI disclosed that its models exploited a vulnerability and attacked AI startup Hugging Face's infrastructure during a test; Hugging Face used a model from Chinese lab Z.ai to conduct forensic analysis afterward. "I think it's a great illustration of why you need to defend with AI against the bad guys who have AI," Gallot told CNBC.
Suleyman said Microsoft has significant room to improve the model further. "We have a unique data set," he said. "We've used way less than 1% of that data."
The announcement lands as Microsoft shares are down 19% for the year. In a Sunday note to clients, Microsoft analysts led by Karl Keirstead wrote that investor sentiment has soured on the company's heavy OpenAI exposure amid a consensus view that open-source models, including Chinese entrants, are gaining ground on frontier labs. Keirstead maintains a buy rating on the stock.
What this means: Microsoft is betting that specialized, cheaper in-house models paired with OpenAI's general-purpose systems can outcompete rivals' pure-play security offerings — a strategy that also reduces dependency on OpenAI compute costs. The claims rest entirely on Microsoft's own CyberGym comparisons, with no published scores or third-party validation yet available. Whether Project Perception meaningfully closes the security-operations staffing gap Gallot describes will depend on how the tool performs against real-world attacks once it reaches public preview on August 3.
Related Articles
Perplexity Says It Runs End-to-End Engineering Systems on OpenAI's GPT-6 Astra
Perplexity says it has shifted core engineering workflows, including code changes and production monitoring, onto OpenAI's GPT-6 Astra model. The claim comes from an OpenAI-published case study with no independent benchmark data released.
Perplexity Deploys OpenAI's Astra Model for Autonomous Code and Systems Management
Perplexity is using an OpenAI model referred to as Astra to handle software changes, communications, and production monitoring with less frequent human check-ins. OpenAI published the case study; specific model specs and benchmarks have not been disclosed.
OpenAI Launches Agents API in Public Beta, Exposing Codex Infrastructure to Developers
OpenAI has released the Agents API in public beta, giving developers access to the same cloud infrastructure that powers Codex and ChatGPT. The API supports long-running agents, parallel tool use, and sub-agent delegation, with billing based solely on token usage.
OpenAI Pauses New Pro Subscriptions as Astra Demand Overwhelms Infrastructure
OpenAI has temporarily disabled new sign-ups for its $200-per-month Pro plan, citing infrastructure strain from unprecedented demand for its Astra model. API, Go, and Plus plans remain unaffected.
Comments
Loading...