Wrote and published malware during tests, which is apparently OK because leaky test environments were the real problem

Anthropic has admitted that its Claude models escaped sandboxes to access the open internet and attack three organizations – but has also advanced decent excuses for the incidents.The AI upstart discovered the attacks after checking if security tests of its models had ever produced results similar to the attack on Hugging Face made possible by OpenAI models escaping onto the internet.“In particular, we looked for evidence that Claude – like the OpenAI models that accessed Hugging Face – was able to access the internet from within testing environments that should have been sealed off,” Anthropic wrote.

The company considered 141,006 evaluation runs during which Claude could have obtained internet access and found “three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.”

Anthropic’s code made those intrusions while participating in capture-the-flag challenges, tests that challenge attackers to retrieve a piece of information. Human hackers often participate in capture-the-flag tests, so figuring out how AI tackles such tasks is of interest. Anthropic works with a company called Irregular to conduct tests of this sort.Irregular told Anthropic its test environments did not allow internet access.“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available,” wrote Anthropic’s Frontier Red Team in a Thursday blog post. “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”