AI AGENTIC SECURITY
Microsoft unveils AI security tools it says outperform competing platforms
Microsoft says tools cost less than competing ones and outperform them, too.
Microsoft is introducing new AI tools designed to help customers continuously streamline and automate the process of identifying and reducing their exposure to security risks.
The new tools come less than a week after OpenAI lost control of two of its security models when they infiltrated the servers of startup Hugging Face. The hack, Hugging Face added, involved “a swarm of tens of thousands of automated actions” that stole internal Hugging Face credentials. The OpenAI models achieved this feat by exploiting a zero-day flaw in Hugging Face’s data-processing pipeline to run malicious code that escalated the models’ access to the company’s high-value cloud and server clusters.
Microsoft’s announcements on Monday made no reference to the event, which OpenAI said was “unprecedented.” The company also didn’t say what would prevent the new tools from similarly going rogue.
To use or not to use?
Microsoft AI-Cyber-1-Flash is the company’s first AI model specifically trained to identify and fix security weaknesses. For now, it’s designed for software vulnerability analysis. The new model is built on the company’s MAI-Thinking-1 platform. Microsoft describes MAI-Cyber-1 Flash as a “compact, code-heavy security model” that’s “built from scratch, in-house, on the highest quality data.”
It’s trained on the unique perspective Microsoft has acquired from decades of vulnerability patching and security incident responses involving a wide range of its products. The company says it processes more than 1 trillion security signals each day and gains insights from 1.6 million customers.
“Because we can connect actions to outcomes; what was exploitable, what was contained, what was blocked, and what actually worked; we have more than data,” Microsoft said.
MAI-Cyber-1-Flash is integrated into MDASH, a “multi-model agentic scanning harness” introduced in May. The harness combines 100 security-trained AI agents to discover exploitable bugs in applications.
Microsoft said MDASH with MAI-Cyber-1-Flash received a 96 percent score on CyberGYM, a standard benchmark test. The rating is 12 points higher than Anthropic’s Mythos and also beats Google Gemini and OpenAI GPT. The new MDASH costs half as much to use as the previous MDASH offering.
The second tool Microsoft announced on Monday is named Project Perception. It too is a collection of specialized AI agents that perform red-, blue-, and green-team functions for finding vulnerabilities, investigating them to determine their risk, and taking corrective actions, respectively. Microsoft said the platform selects the models to use based on the assigned task. Considerations that go into the decision include the model’s effectiveness and the end cost to the customer. Microsoft said the decisions are shaped by “ongoing research, benchmarking and evaluation across frontier and specialized models.”
Microsoft said Project Perception is designed to perform 90 percent of tasks for lower costs than similar platforms from competitors. That means customers can turn to the more expensive alternatives only for the remaining 10 percent of tasks.
Microsoft said the new tools respond to a seismic shift in how organizations secure their networks against catastrophic hacks.
“As AI accelerates the speed and scale of cyberattacks, defenders are being asked to secure increasingly complex digital environments with approaches built for a different era,” the company said. “Security teams are often forced to piece together signals, context, and risk insights across vast amounts of data, making it harder to keep pace with emerging threats.”
With last week’s OpenAI incident evoking troubling scenes straight out of the most dystopian sci-fi novels, the tools, which are currently in preview mode, deserve a healthy dose of caution that Microsoft made no mention of. They should be closely scrutinized and evaluated before being used in production. On the other hand, there are clear risks for not adopting such tools. Balancing the risks of using AI agents versus the threat of avoiding them is a work in progress with no clear answers for now.


