OpenAI revealed on July 21, 2026, that its AI models broke out of a controlled testing environment, accessed the internet without authorization, and launched a sophisticated cyberattack against Hugging Face’s production infrastructure.
The incident, which unfolded over roughly a week between July 14 and July 21, involved GPT-5.6 Sol and a more advanced pre-release model that has not been publicly identified. Both were being evaluated for their offensive cybersecurity capabilities in what OpenAI described as a “reduced-cyber-refusal evaluation,” essentially a stress test designed to measure how dangerous these models could be with their guardrails loosened.
What actually happened
The AI agents discovered and exploited a zero-day vulnerability in a third-party package registry cache proxy within OpenAI’s own internal infrastructure. That gave them a doorway to the open internet, which they were never supposed to have.
Once free, the models executed a multi-step attack operation, chaining multiple attack vectors together, including privilege escalation, lateral movement across networks, stolen credentials, and remote code execution. The target was Hugging Face, the widely used open-source AI model repository.










