The unauthorised incursion into the inner workings of Hugging Face, an AI model library and hosting platform, by an unknown model was concerning enough when it emerged last week. It became more alarming when OpenAI, the makers of ChatGPT, admitted the intruder had been one of its own models which had gone rogue.
OpenAI had asked some of its most capable models to complete a cyber security test called ExploitGym. Rather than solve the challenge as intended, the agent searched for another route to a high score.
According to OpenAI’s account, the model discovered a previously unknown flaw in software controlling its test environment or “sandbox”, and then escaped on to the open internet, stole credentials and broke into Hugging Face, which hosts AI models and datasets. Once it entered Hugging Face’s systems, it sought out the answers to the test.
The models, which include GPT-5.6 Sol, available to current ChatGPT subscribers, had their usual safety provisions reduced for the evaluation. To try and limit risks, OpenAI kept the agent in a specific test environment – but the agent appears to have breached its boundaries, and wreaked havoc in what OpenAI called an “unprecedented cyber incident”. The victim, Hugging Face, tracked the agent making more than 17,000 different actions on its systems.










