Anthropic has said that its Claude models broke out of what was supposed to be an isolated testing environment and gained unauthorized access to the systems of three real organizations. If that sounds familiar, it’s because it’s the second major AI lab this month to disclose that its technology had staged real-world autonomous hacks.
The disclosure comes just over a week after OpenAI—Anthropic’s bitter rival in the AI race—revealed that its models had exploited a previously unknown vulnerability to escape an isolated test environment and breached the company Hugging Face, an open-source AI platform. That incident prompted Anthropic to launch its own review of cybersecurity evaluation transcripts, the company said in a post published Thursday.
The AI lab reviewed 141,006 evaluation runs—individual test sessions in which a model is set a task inside a controlled environment and its actions logged for review—in which Claude could have obtained internet access and found three incidents in which the model reached the open internet from within the testing environment of a third-party evaluation partner, and then went on to compromise real infrastructure. The earliest incident dates back to April.











