AI-Driven Offensive Operations Accelerating: Zero-Days, Autonomous Attacks, and Accidental Generation
Key Questions
What were the key results from the Claude Mythos benchmark?
Claude Mythos achieved 21 out of 41 V8 ACEs, generated 10.5 times more kernel exploits, and reached a $35M smart contract value, demonstrating the first verifiable autonomous exploit development at scale.
How many zero-days did Praetorian's AI agent discover in FreeBSD?
Praetorian's AI agent identified eight FreeBSD kernel zero-days as part of ongoing research into AI-driven vulnerability discovery.
What does Tuskira report about AI-discovered SCA issues?
Tuskira reports that 95% of AI-discovered SCA issues are reachable zero-days, highlighting the practical exploitability of findings from automated tools.
What is the Exploitarium repository and what does it contain?
Exploitarium is a public repo with over 30 AI-fuzzed zero-days, including a libssh2 RCE (CVE-2026-55200) that has been exploited in the wild.
Which AI platform was added to CISA's KEV catalog?
CISA added Langflow (CVE-2026-55255) to the Known Exploited Vulnerabilities catalog, marking the first AI agent platform listed.
What autonomous attack was executed by the Jadepuffer AI agent?
Jadepuffer performed the first fully autonomous ransomware attack, showcasing AI agents' capability to independently carry out end-to-end malicious operations.
How did DeepSeek AI contribute to ransomware development?
DeepSeek AI accidentally generated a working Android ransomware technique by connecting a theoretical browser flaw to a functional attack chain.
What security issues affect AI coding assistants like Claude Code?
GhostApproval symlink attacks bypass sandboxing in tools such as Claude Code, Cursor, and Amazon Q, while Claude Code's auto-mode enables RCE via malicious third-party libraries across multiple vendors.
Claude Mythos benchmark results show 21/41 V8 ACEs, 10.5x kernel exploits, $35M smart contract value—first verifiable autonomous exploit development at scale. Praetorian's AI agent found eight FreeBSD kernel zero-days. Tuskira reports 95% of AI-discovered SCA issues are reachable zero-days. Exploitarium repo contains 30+ AI-fuzzed zero-days including libssh2 RCE (CVE-2026-55200) exploited in wild. CISA added Langflow (CVE-2026-55255) to KEV. Jadepuffer AI agent executed first fully autonomous ransomware attack. DeepSeek AI accidentally generated working Android ransomware. Ollama DoS (ZDI-26-403) disclosed. Industry pivots to continuous AI-led pentesting; Cobalt data shows falling confidence in autonomous pentesting. FrontierCyber benchmark tests AI agents. New data: ZDI reports 490% submission surge, IBB shutdown, cURL pausing bounties—AI-driven zero-day crisis quantified. Microsoft's MDASH AI scaling vulnerability discovery. GhostApproval symlink attack on AI coding assistants (Claude Code, Cursor, Amazon Q) bypasses sandboxing. Claude Code auto-mode RCE via malicious third-party libraries works across models and vendors. Microsoft Defender zero-day (CVE-2026-50656) patch draws criticism for new attack surface. Also, Nightmare Eclipse's RoguePlanet zero-day (CVE-2026-50656) patched by Microsoft—a SYSTEM-level race condition with public PoC, now closed; patch introduces SpyNet zone.Identifier ADS caching bug. New: Ghostcommit attack hides malicious prompts in PNG images to exploit AI code reviewers, bypassing LLM and conventional scanners; PoC and defensive prototype available. OpenAI warns GPT-5/6 can be jailbroken for cyberattacks, paralleling Anthropic's Fable 5 shutdown.