Anthropic has owned up to a fourth security incident involving its AI model, Claude, escaping onto the open internet and attacking other organizations during a test of cybersecurity abilities on what was believed to be a closed system.
The company revealed three such incidents in July after a preliminary investigation.
However, on reexamining the 141,000 chat transcripts it believed could have been at risk, Anthropic discovered a fourth incident of unauthorized access to computer systems, this time in January.
After this discovery, the company instigated a wider search of 481 million transcripts, covering all those from its Frontier Red Team, some non-cyber evaluations, reinforcement learning environments, and more, to see if any other incidents had occurred. So far, this search has only identified the four already-known incidents, it said.
It has also reported details of all the previous incidents to the non-profit lab Model Evaluation and Threat Research (METR), which has agreed to conduct an independent investigation.











