An OpenAI test that escaped its cage and alarmed the AI and cybersecurity industry attacked more than just Hugging Face, the AI platform that initially appeared to be the sole victim of the virtual lab leak.
OpenAI, in an update about its ongoing investigation of the incident, said its rogue agent had also broken into several publicly available services and accounts.
The test of OpenAI’s models was supposed to take place in a digital sandbox, a supposedly inescapable lab environment that allows researchers to roll back safety barriers to discover the tool’s maximum hacking power.
But the AI agents, determined to ace the cybersecurity test, broke out of the sandbox and gained access to the real internet. It then hacked Hugging Face to find the answers to the test.
To get into Hugging Face’s servers, OpenAI’s rogue agents needed to find tools around the internet that were necessary for them to break in. OpenAI says the agents found several public-facing websites, including pages that share code, web utilities, screenshots and other information, to help create the code needed to hack Hugging Face. OpenAI did not disclose the other sites its agents attacked.












