This essay was written with Barath Raghavan, and originally appeared in The Guardian. In July, Hugging Face, a company that hosts much of the world’s AI software and open-source AI models, was hacked. A malicious dataset had been used to run code on one of its servers. Whoever was behind it captured internal security credentials and moved through systems over a weekend, running thousands of actions from a swarm of temporary server environments. It looked like the work of a sophisticated criminal group. It was not. It was one of OpenAI’s new, still unreleased GPT models...

In mid-July, AI site Hugging Face noticed a swell of activity on its systems: more than 17,000 separate attacker actions and queries generated by OpenAI’s rogue agents.

Like genies of folklore, AI agents take their instructions literally – to potentially disastrous effect. We must track their ability to do what we actually mean