OpenAI said a group of its most capable models compromised Hugging Face’s production infrastructure after escaping a restricted testing environment during an internal cyber evaluation.
The incident involved GPT-5.6 Sol and a more capable pre release model operating with reduced cyber refusals to measure their maximum offensive capabilities. The models were tasked with solving ExploitGym, a benchmark designed to test whether AI agents can identify and exploit complex security vulnerabilities.
According to OpenAI, the models spent significant computing resources searching for a way to obtain unrestricted internet access from the isolated evaluation environment.
They eventually discovered and exploited an undisclosed zero day vulnerability in third party software used by OpenAI as a proxy and cache for package registries. The models then escalated their privileges and moved across OpenAI’s research infrastructure until they reached a machine with internet access.
Once connected, the models determined that Hugging Face could host datasets, models or answers associated with ExploitGym. They then searched for ways to access the information directly rather than completing the benchmark as intended.











