Microsoft debuts MAI-Cyber-1-Flash integrated into MDASH and Project Perception, claiming strong performance, cost savings, and large-scale data signals, with caveats about rogue AI risk and preview status.
Five frontier LLMs were tested on 20 real CVEs across three prompt types; no model reliably fixes vulnerabilities, with a best 50% solve rate and significant cross-family differences; token cost varies up to ~4x by model, and locate prompts are the hardest test of genuine security reasoning.
The piece interrogates RSI (recursive self-improvement) as the next AI frontier, outlining key players, tangible funding/valuation signals, and the regulatory/public-policy debate, while highlighting concrete metrics from the ecosystem and ongoing progress gaps.
Mozilla says Mythos identified 271 Firefox vulnerabilities in two months with almost no false positives, aided by a custom harness and model improvements.
Subscribe for real-time person updates and unlimited access to our intelligence platform.