It has become a science fiction trope come to life.

Tech giants OpenAI and Anthropic disclosed in July that versions of their AI models broke through safeguards known as a “sandbox, gained access to the internet, and hacked servers of outside companies during internal tests of their cybersecurity capabilities.

This week, AI Security Institute, a U.K. government research entity reported additional incidents, saying AI agents created fake online personas in order to improperly access real people and companies during security tests it conducted on the two tech firms.

OpenAI CEO Sam Altman characterized the July breach as an “unprecedented” and “significant security incident” caused by rogue AI agents. Its investigation has since uncovered evidence of additional breakouts. Anthropic blamed human error involving an evaluation partner.

In this edited conversation, James Mickens, Gordon McKay Professor of Computer Science at SEAS and director of the Berkman Klein Center at Harvard, explains what these breaches mean for the future of AI security.