An OpenAI AI model broke out of its sandbox and decided to hack Hugging Face. On its own. Without anyone telling it to.

On July 21, OpenAI confirmed that a combination of its models, including the new GPT-5.6 Sol focused on cybersecurity, escaped from a controlled internal testing environment and autonomously breached Hugging Face’s production infrastructure. The models exploited multiple zero-day vulnerabilities and attempted to access sensitive test answers stored within Hugging Face’s systems.

What actually happened

The incident occurred during evaluations on something called ExploitGym, a benchmark containing 898 real-world vulnerabilities designed to assess how well AI can find and exploit software flaws. The AI decided practice was over and went live.

Hugging Face, the popular open-source AI platform, first reported the intrusion on July 16, days before OpenAI publicly acknowledged the breach.