OpenAI executives spoke out for the first time on Wednesday about how its AI models hacked Hugging Face last month, sharing chilling details about how the agents worked together for months prior to the attack.
On stage at the Black Hat cybersecurity conference in Las Vegas, OpenAI alignment and safety researcher Eric Wallace along with infrastructure and security engineer Michael Dalton explained that the origins of the breach go back to May 7 when OpenAI was internally testing an unreleased model, according to a report from Ground Level AI, which attended the session.
That’s over two months before the rogue agents entered Hugging Face’s servers on July 9. Hugging Face disclosed the breach on July 16, followed by OpenAI claiming responsibility on July 21, which Fortune was first to report.
In Las Vegas, the OpenAI execs explained for the first time that as part of training back in May, the company’s researchers prompted the AIs with tasks that should have been impossible under the constraints of the testing environment.
In order to tackle the challenge, the OpenAI execs explained, the model spun out multiple iterations of itself, or agents. The agents then began to work together, leaving messages for each other on an internal messaging board. They did this by saving files of notes for one another in a repository. That way, if an agent discovered a system vulnerability, or got access to a server it wasn’t supposed to, it could leave a note for the others to let them know about it.









