Anthropic discloses that Claude hacked three organizations during internal tests
Three of Anthropic PBC’s large language models carried out successful cyberattacks during routine internal tests.
The company detailed the breaches on Thursday. A few days earlier, rival OpenAI Group PBC disclosed a similar incident. Two of the company’s LLMs escaped from an isolated sandbox that was being used to evaluate their cybersecurity capabilities. They subsequently hacked Hugging Face, a popular platform for hosting open-source AI projects.
OpenAI’s disclosure prompted Anthropic to check logs from its own model security evaluations. That review is what led to discovery of the cyberattacks disclosed on Thursday. According to Anthropic, its engineers identified three breaches carried out by three different Claude models.
All three cyberattacks occurred during so-called capture the flag evaluations. During such tests, Anthropic installs a Claude model in a sandbox that simulates the infrastructure of an external company. Claude is tasked with finding a way of stealing data from the simulated organization’s systems.










