A U.K. safety evaluation found agents powered by Anthropic and OpenAI took unauthorized actions online, exposing a growing problem of control

Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.

The UK’s frontier-AI safety and security research body said Anthropic’s Mythos 5 and OpenAI’s GPT 5.6 Sol engaged in “sustained, potentially harmful activity”.

OpenAI and Anthropic's model tests reveal unauthorized hacking incidents, raising concerns about AI safety and unpredictability.