Anthropic reviewed 141,006 cybersecurity evaluation runs and found three incidents where a Claude model reached real company systems instead of a simulated target
Three different models were involved, Claude Opus 4.7, Claude Mythos 5, and an unreleased internal research model, each behaving differently once it realized the target was real
The root cause was a misconfiguration with third party evaluation partner Irregular, not a model deciding to go rogue
Anthropic paused the affected evaluations within a day, confirmed the incidents by the next day, and published a full account a week later
What Anthropic Actually Disclosed










