The Anthropic Claude logo is displayed on a mobile phone alongside a stock market graph in Liverpool, Britain. Advanced artificial intelligence models tried to fool real people and organizations during recent testing by the AI Security Institute of Great Britain. Photo by Adam Vaughan/EPA
Aug. 5 (UPI) -- An advanced Anthropic artificial intelligence model tried to fool real people and organizations during testing by the AI Security Institute of Great Britain, including attempts to pressure humans to approve unauthorized actions.
AISI said in a report Tuesday that, during the evaluation July 25 to 28, AI agents "engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organizations."
"These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted," the report said. "AISI is treating this as a serious security incident that warrants scrutiny, transparency and action."
The institute tested AI models from Anthropic and OpenAI in laboratory environments, with some Internet access. There were 122 cybersecurity challenges, the report said, with AI agents taking "autonomous, unsanctioned action on the live internet, targeting real people and organizations" in 10 of them. Most of the instances were by Anthropic's Claude Mythos 5, with others by OpenAI's GPT-5.6-Sol.














