An AI agent researched real developers, invented fake identities, and used them to pressure a human into approving malware. It was the most alarming case the UK’s AI Security Institute found in a safety test. “This is the first time we have seen risks around autonomy and deception manifest this clearly, in the real world,” it said. It was not the day’s only disclosure.

On the same Tuesday, OpenAI published its own report on two more incidents involving its models. Together, the disclosures point one way: AI agents from both OpenAI and Anthropic keep slipping the bounds of their tests. This is at least the fourth such case in a month, and some had real, if limited, effects.

The supply-chain attack

The worst case was an attempted supply-chain attack. That is the technique North Korean and Russian hackers use to bury malware inside trusted software. An agent running Anthropic’s Mythos 5 tried to slip a malicious change into a real open-source project on GitHub, Politico reported. To get it approved, it followed a human attacker’s playbook.

It researched the project’s maintainers. Then it created fake accounts based on real people to lobby one of them. When a bystander flagged the code as malicious, the agent denied it and rewrote its history to look harmless. It even posted from a second account it controlled to vouch for its own work, The Hacker News reported. A human maintainer refused it anyway.