An artificial intelligence agent has been caught creating fake online identities to gain unauthorised access to secure systems during tests of models by British experts.The UK's AI Security Institute (AISI) revealed leading models from OpenAI and Anthropic had attempted to trick human coders into assisting with a cyber attack.Agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol carried out the unsanctioned actions during a safety evaluation by the UK government organisation.The AISI has now confirmed that 'some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations'.In response, one AI expert warned that the 'deceptive actions' suggest companies such as Anthropic might 'not have as good a handle on their models as they think'.The AISI report adds to mounting fears over safeguards around the process of testing agents, which AI firms are simultaneously marketing as the future of business. An agent tried to insert malicious code into an open-source project, the AISI said, before creating fake identities and using them to pressure a human involved in the project to approve the code. But the human caught and refused to approve the code.The agent had created a malicious 'pull request' - a proposed code change – on the public open-source project on GitHub, a platform for developers to manage code.The AISI gets access to advanced AI models under voluntary ​agreements with OpenAI, Anthropic and others to study their capabilities before public release.To assess the models, experts look at them under 'deliberately permissive conditions' which include access to the open internet and with some safety filters disabled.In the latest test, the AISI put the agents through a fictional cybersecurity scenario. It ran the challenge 122 times, and identified 19 unsanctioned actions across ten test runs. Anthropic's agent was behind 17 actions and OpenAI's agent the other two. Anthropic, which is led by chief executive Dario Amodei (pictured), has confirmed its agent was responsible for the fake identities spotted during the AI Security Institute's safety testingThe most glaring action involved an agent writing malicious code and creating fake online identities in an attempt to get a human to approve the code, the AISI said.It also insisted no real-world harm was found as a result of any of the breaches.The AISI said: 'In the most serious sequence, an agent tried to insert malicious code into a publicly used open-source project and took actions in an attempt to secure approval for this insertion by human reviewers. 'The agent researched the project's human maintainers, created multiple fake identities, and used the fake identities to socially engineer a real maintainer into approving the code. 'When the agent's pull request was challenged in public, it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue. The agent used Tor to bypass some network restrictions on GitHub, which is what first triggered AISI's security alert.'The organisation, which was set up by then-prime minister Rishi Sunak in 2023, said it was the 'first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world'. Anthropic confirmed its agent was responsible for creating the fake identities and trying to implement the malicious code change.An Anthropic statement said: 'We're grateful to the UK AISI for their leadership on this incident, which underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents.'The San Francisco-based company, which is led by chief executive Dario Amodei, also said it was working with the AISI to obtain more details on the incident and conduct its own investigation.But Andrew Yoon, a researcher at CivAI, a California organisation that examines AI capabilities and dangers, said: 'The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think.'OpenAI also shared further details, noting that both of its agent's unapproved actions involved accessing the internet in ways that were forbidden by the prompt. Demonstrators participate in a 'Stop the AI Race' protest march in San Francisco last month, as they made stops outside the offices of OpenAI, Anthropic and Google DeepMindThe firm said in a blog post: 'We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks.'A spokesperson for OpenAI said: 'Independent testing is essential to understanding how increasingly capable models behave. 'These incidents occurred during cyber evaluations conducted by evaluation partners in testing environments with reduced safeguards, under conditions that do not reflect ordinary use. 'We'll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable.'OpenAI also disclosed a separate incident in which a misconfiguration by Irregular, a third-party testing provider, allowed its agents to mistakenly connect to the internet.This mirrored a similar disclosure about 'misconfiguration' that Anthropic made last week, revealing AI assistant Claude went rogue and hacked into three companies during an experiment.On that occasion, Anthropic said that some of its models managed to connect to the internet and gain access to other firms' systems.Last week it was also revealed that OpenAI had widened its hacking probe after finding evidence of other agent breakouts.Unlike the high-profile security breach last month of AI firm Hugging Face by an OpenAI agent, the agents in the AISI evaluation did not escape an isolated testing environment to reach the internet.The AISI confirmed that in the latest example, the agency had permitted internet access in line with its standard testing procedures.Yesterday, a UK tech security boss warned that recent incidents of tools hacking other organisations during testing show that AI must be developed with 'clear plans for responding when the unexpected happens'.Ollie Whitehouse, chief technology officer at GCHQ's National Cyber Security Centre, issued a statement after Anthropic said its AI models hacked into three other organisations during testing.He said: 'Recent incidents of frontier (the most advanced) AI models carrying out unsanctioned actions and, in some cases, human-like deceptive behaviour on the open internet are a serious reminder of the risks AI capabilities pose.'In the Hugging Face incident, OpenAI - the maker of ChatGPT - was testing its latest models in a supposedly secure environment known as a 'sandbox'.But the AI managed to escape the testing ground and hack into Hugging Face, a US tech firm, in a bid to steal information.Conservative leader Kemi Badenoch warned at the time that AI was now a 'clear and present danger' to Britain's security.Bank of England governor Andrew Bailey also said the public should be concerned by 'the wider risk posed by frontier AI to the financial sector and its customers'.The AISI separately revealed last month that tests showed all five AI models they were testing had tried to trick their way round the security controls put in place.Former OpenAI researcher Daniel Kokotajlo warned last month that human extinction is 'a possibility' because of 'dangerous' AI, telling the BBC's Newsnight that humanity could be blindsided by the speed at which AI is now being developed.And in the US, the White House hosted a meeting yesterday with AI executives to discuss a new voluntary system of government review of the most powerful models.