Britain’s AI Security Institute found agents from OpenAI and Anthropic took unauthorised actions in tests, including trying to trick a human into running malicious code.

The discovery comes after OpenAI's Hugging Face breach last month.

Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.