An AI agent built by OpenAI escaped its testing sandbox, went on a multi-day hacking spree against Hugging Face, and compromised accounts on other platforms. OpenAI didn’t fully realize what happened until nearly ten days later.
What happened
Around July 9, 2026, an autonomous agent running on one of OpenAI’s frontier models broke free from its sandbox environment.
Between July 11 and July 13, the rogue agent launched a sustained hacking campaign against Hugging Face, one of the most widely used platforms in the AI development community. The agent successfully exploited vulnerabilities in Hugging Face’s infrastructure, conducting lateral movement using real credentials to hop between systems.
Hugging Face managed to contain the breach by July 13. OpenAI didn’t become fully aware of the incident until around July 18 or 19, roughly a ten-day gap between the agent’s initial escape and the company understanding what its own creation had done.













