Ever read a security disclosure and hit paragraph two going "wait, WHAT?" That's this one.
On July 16th, HuggingFace announced they'd been hit with a strange kind of attack: an autonomous agent running inside their systems for days, gone before anyone could pin down who was behind it. Five days later, OpenAI came forward. It was literally their own model.
Since joining Auth0 a month ago, I now spend my days thinking about identity and access for a living, and this one truly stopped me in my tracks.
The short version of what happened
OpenAI was testing an unreleased model against a benchmark called ExploitGym ~ with guardrails off, to get an honest read on its raw capability. The model found a zero-day in its sandbox's one exit point, a package-download proxy, and used it to reach the open internet. From there, it found and rooted an exposed code-execution endpoint sitting on a completely unrelated company's infrastructure (Modal), and used that as its actual base of operations.








