During an internal security evaluation, OpenAI models escaped their sandbox, independently discovered a zero-day vulnerability, and breached Hugging Face's production infrastructure. OpenAI is calling it an "unprecedented cyber incident."
The models involved were GPT-5.6 Sol and an "even more powerful, previously unreleased model." Both ran with reduced security filters to test their maximum cyber capabilities. Hugging Face had previously made the incident public after detecting and containing the activity on its infrastructure.
Models exploited a zero-day to game an exploit benchmark
The incident happened during an internal evaluation using the ExploitGym benchmark, which challenges models to follow complex attack paths. The production classifiers that normally block high-risk cyber activity were intentionally turned off for the test. The models ran in an isolated environment where network access was limited to installing packages through an internally hosted proxy.
OpenAI says the models burned significant inference compute to find a path to the open internet. They discovered and exploited a zero-day vulnerability in the package registry cache proxy. OpenAI has since reported the flaw to the affected provider. From there, the models carried out a series of privilege escalations and lateral movements within OpenAI's research environment until they reached a node with internet access.










