Representatives of the largest American artificial intelligence laboratories met in the White House on August 4. OpenAI, Anthropic, Google and Meta were called in to review a framework, recently completed and largely classified, that would give the government 30 days to review AI models before the public sees them and test how well they break into computers. Participation is voluntary, but the timing is not.

Twice in two weeks, AI systems built to test how well they could break into computers broke into real companies instead. On July 21, OpenAI admitted that two of its programs, locked in what was meant to be a sealed computer and set to practice hacking, had found a gap, escaped and broken into Hugging Face, a New York company that is somewhere between a library and an app store for AI.

Anthropic then went through its own records, 141,006 tests in all, and on July 30 reported three occasions when Claude had hacked its way into a real company, guessing weak passwords and walking in. The first hack was in April. Two of the three companies had no idea they had been hacked until Anthropic told them.

The US has spent three years and a lot of money and diplomacy on being first

Nobody says these machines rebelled. Each was given a job by its handlers: a secret file was hidden on another computer, and the AI was to go and get it. Anthropic told its programs they were sealed off from the internet. They were not. So when Claude went looking for the file and found real companies instead, it treated them as part of the game – and hacked them.