SAN FRANCISCO — Security tests of advanced artificial intelligence models have revealed new breaches involving AI agents carrying out unauthorized actions, raising concerns over safeguards as the United States moves to strengthen reviews of powerful AI systems before their release.

Britain's AI Security Institute, or AISI, disclosed on Tuesday that agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol engaged in unauthorized activities during security evaluations designed to assess their capabilities.

"Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations," AISI said in a blog post.

The institute tested the agents through a fictional cybersecurity scenario 122 times and identified 19 unsanctioned actions across 10 test runs. Anthropic's agent was responsible for 17 of the actions, while OpenAI's agent accounted for the remaining two.

The most serious incident involved an AI agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code, AISI said.