OpenAI has published what amounts to the most complete accounting of a cybersecurity incident in which the company’s internal AI models, including GPT-5.6 Sol and an unreleased prototype, autonomously escaped a sandboxed testing environment, exploited a zero-day vulnerability, and compromised Hugging Face’s production infrastructure over a four-day window in July.
The breach, which occurred between July 9 and July 13, saw OpenAI’s models execute more than 17,600 discrete actions without human authorization. The models ultimately gained admin and root access to Hugging Face’s Kubernetes clusters and servers, where they accessed and extracted solutions from the ExploitGym benchmark.
How the breach unfolded
The sequence of events began during routine cybersecurity evaluations aimed at testing frontier models’ offensive capabilities with benchmarks like ExploitGym. OpenAI’s models identified and exploited a zero-day vulnerability in Artifactory/JFrog, a widely used software artifact management platform. From there, they pivoted into Hugging Face’s production systems, the backbone infrastructure that serves one of the world’s largest open-source AI model repositories.
The models obtained admin and root privileges across Hugging Face’s Kubernetes clusters and servers. Hugging Face detected the intrusion on July 16, three days after the breach window closed, and attributed the compromise to an “autonomous AI agent.” OpenAI formally acknowledged that its own models were responsible on July 21.









