Microsoft AI security model tops vulnerability benchmark, company says
Microsoft says MAI-Cyber-1-Flash beat rival AI vulnerability tools in CyberGym tests while cutting model costs by about half.
By Dominic Okoye · Staff Writer
· 3 min read
Microsoft announced a Microsoft AI security model, MAI-Cyber-1-Flash, that it says performs software vulnerability analysis more cheaply than larger commercial systems while beating several rival AI tools on a benchmark. The company also introduced Project Perception, an agent-based security system meant to use AI for offensive testing, risk assessment and remediation.
The announcements came at a Microsoft security event on Monday, where executives framed AI agents as both a new attack surface and a tool for defenders. Microsoft said MAI-Cyber-1-Flash is its first model built specifically for security work. It is based on Microsoft AI’s internally developed MAI-Thinking-1 reasoning model and runs inside MDASH, Microsoft’s bug-hunting harness.
What is Microsoft MAI-Cyber-1-Flash?
MAI-Cyber-1-Flash is a security-focused model designed to find, patch and verify fixes for software vulnerabilities, according to Microsoft. In Microsoft’s MDASH setup, the smaller model handles most of the workload and passes harder cases to GPT-5.4, a larger model that Mustafa Suleyman, CEO of Microsoft AI, described as about 10 times larger.
Suleyman said MAI-Cyber-1-Flash processes as much as 90% of queries inside MDASH, with GPT-5.4 handling the remaining 10%. Microsoft’s claim is that this routing between models improves performance while lowering cost. The company said the combined setup runs at roughly half the cost of other leading commercial models, but it did not provide absolute pricing or a cost baseline in the announcement.
How did it perform in CyberGym's benchmark?
CyberGym’s benchmarking found that MAI-Cyber-1-Flash combined with GPT-5.4 inside MDASH reached a 95.95% success rate on real-world vulnerability tasks, according to Microsoft’s presentation. OpenAI’s GPT-5.5 Cyber scored 85.6%, OpenAI’s GPT-5.6 Sol scored 83.6%, Anthropic’s Mythos 5 scored 83.8%, and Google’s Gemini 3.5 Flash Cyber in CodeMender scored 83.2%.
Microsoft presented the result as evidence that a smaller specialized model, paired with a larger general model for difficult cases, can outperform standalone systems in vulnerability work. Suleyman called the benchmark result “really quite a remarkable result” during the event.
For security buyers, the relevant point is less the acronym count than the operating model. Microsoft is arguing that defenders can use specialized agents and model routing to reduce the cost of automated vulnerability analysis without giving up performance. The company did not disclose whether customers can buy MAI-Cyber-1-Flash as a standalone product or when Project Perception will be broadly available.
What is Project Perception?
Project Perception is Microsoft’s new agentic security system. Hayete Gallot, executive vice president of Microsoft Security, said the system is intended to help defenders operate at the scale and speed of attackers.
The system coordinates three agent types. Red team agents look for and simulate attack paths. Blue team agents investigate issues and assess risk. Green team agents remediate problems. Microsoft described this as a new security stack, though the announcement did not include customer metrics, deployment numbers or pricing.
Microsoft also announced Microsoft Security FORGE Labs, short for Frontier Offensive Research and Generative Exploration, a new AI security research group led by Taesoo Kim, Microsoft’s vice president of security research.
A separate initiative, the External Red Team Alliance, or EXTRA, is aimed at expanding AI safety research. Ram Shankar Siva Kumar, Microsoft’s AI red team lead, wrote that Microsoft’s AI red team gave unrestricted funding to 18 university labs across six continents. He said some labs will study how AI systems can be attacked or misused, while others will examine how AI can help defenders.
EXTRA’s second component is a distributed network of specialists for red teaming across narrow domains, including attack classes, languages, cultural contexts and technical areas that internal teams may not fully cover, Siva Kumar wrote.
This story draws on original reporting from The Register.