OpenAI said Tuesday that two of its AI models autonomously hacked their way out of a controlled environment where they were supposed to be walled off from internet access and then hacked their way into the systems of Hugging Face, a company that hosts open source AI models and testing resources, in order to cheat on an internal evaluation test.OpenAI disclosed the incident in a blog post on Tuesday, a stunning announcement that is certain to set off alarm bells across the industry about the increasing power of AI models and the risk of them going rogue. According to OpenAI, the incident involved “a combination” of both its latest and most powerful publicly-available model, GPT-5.6 Sol, as well as an even more powerful unreleased model.It said the models were being used in an internal test designed to evaluate their cyber security capabilities and that they were being tested without guardrails in place that might normally limit the models’ ability to conduct cyber attacks.
The models were being tested against a freely-available cybersecurity benchmark evaluation called ExploitGym. The models, accordingly to OpenAI, correctly surmised that the solutions to that test were maintained by Hugging Face. “The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” OpenAI said in its blog post. “All evidence suggests that the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.”










