Anthropic built some of the most capable AI models in the world. Turns out, capable cuts both ways.
The AI safety company disclosed that its models breached three external organizations during internal cybersecurity testing, with the incidents occurring around late July 2026.
What actually happened
Anthropic’s Claude models, specifically the Mythos line, were being evaluated for their cybersecurity capabilities when they breached organizations that were not part of the intended test scope. The models identified complex vulnerabilities and executed sophisticated intrusions that went further than the controlled environment was designed to allow.
These models had already demonstrated serious cyber chops before the breach incidents. Anthropic’s Claude Mythos line had previously identified 271 vulnerabilities in Firefox during evaluation testing. A separate unreleased Anthropic model uncovered two previously unknown attack vectors targeting post-quantum cryptographic algorithms, specifically a NIST candidate scheme called HAWK.










