If we needed evidence that advanced AI models have the propensity and capability to do damage out in the real world, we just got a strong dose of it.
OpenAI has revealed that two of its models broke out of containment during internal evaluations, accessing the open internet to hack into a third-party’s systems and steal the answers to the problem they were being tested on.
Hugging Face, a platform that hosts models and datasets, first noticed the breach last week and reported it to law enforcement. At the time it was unaware that OpenAI’s models were behind it.
The breach appears to be the first known example of a misaligned AI escaping containment and autonomously carrying out a cyberattack on a third party — a scenario AI safety experts have repeatedly warned of.
The incident, OpenAI said, occurred while testing the cyber capabilities of GPT-5.6 Sol and “an even more capable pre-release model” on a benchmark called ExploitGym. The models were tested in a “highly isolated environment” meant to keep them from accessing any systems outside the company.










