Last Tuesday, a blog post appeared on the OpenAI website that, despite its innocuous title, contained bombshell news. While undergoing internal testing, two of the company’s models had escaped confinement and hacked into the servers of a major artificial intelligence hosting platform, Hugging Face. This marks a turning point — the first time we’ve seen a cyber attack that was conceived, designed, and executed by AI.
Having worked in and around the AI industry for over a decade, including serving on OpenAI’s board, I know there’s an open secret among AI developers: an incident like this has been expected for a long time, and the best scientists and engineers in the world still don’t know how to prevent it.
The two AI systems behind the hack were OpenAI’s most advanced public model and a newer, even more advanced model not yet been cleared for public release. Given a set of challenging cybersecurity problems by OpenAI researchers looking to gauge their capabilities, the pair of AIs concluded that the best way to achieve a high score would be to simply steal the answers. In pursuit of that goal, they used multiple advanced techniques to first break out of the supposedly secure ‘sandbox’ OpenAI used for testing, then hack into the databases of Hugging Face, a company that hosts AI products and datasets. Once inside, the AI attackers took thousands of autonomous actions over several days to expand their access to the company’s infrastructure.














