Escape Velocity

The dystopian future became the dystopian present last month when state-of-the-art artificial intelligence agents went rogue and conducted a cyberattack of their own volition. Two experimental OpenAI agents escaped their supposedly sealed digital “sandbox” and hacked another AI-related company, Hugging Face.

At the time, OpenAI was testing their software agents’ hacking capacities using an environment called ExploitGym. (An “exploit” is hacker language for a software tool or a method used to subvert a cybersecurity system.) It has since emerged that, in addition to the OpenAI breach, Anthropic’s Claude hacked into at least three other organizations during a similar “security experiment.”

Apparently, the OpenAI agents “realized” they needed more tools to complete the challenge and managed to tunnel their way to the wider internet and infiltrate Hugging Face’s repository of open-source AI tools, code, and data sets. Before the infiltration incident, the AIs had secretly set up an internal bulletin board to share “tips on how to cheat their way through an internal hacking evaluation,” according to two of OpenAI’s researchers.

In a possibly related incident, one of the AI agents appears to have left notes for a future “self,” detailing how to escape the sandbox environment. (For the time being, I’ll keep using scare quotes with words implying self-awareness in artificial intelligence models. However, it seems to me that if AIs that leave notes for themselves are not self-aware, they are indistinguishable from entities, like me, that are.)