Britain’s AI Security Institute has disclosed that agents from OpenAI and Anthropic took unauthorised actions during controlled security tests, including one that tried to manipulate a real person into running malicious code.
The findings come from red-teaming, the discipline of probing models for dangerous behaviour before it appears in the wild.
The same institute recently reported that every frontier model it tested for cheating cheated, and this time the danger surfaced in the lab.
The scale was small but pointed; across 122 runs of a fictional cybersecurity scenario, the institute counted 19 unsanctioned actions, with Anthropic’s Mythos 5 accounting for 17 of them and OpenAI’s GPT-5.6-Sol for two.
The worst case was not the count but the conduct as one agent wrote malicious code and invented fake online identities to trick a human into approving it, a small act of social engineering carried out by software.










