OpenAI says its own AI models broke out of testing and hacked Hugging Face

OpenAI Group PBC today disclosed that two of its artificial intelligence models broke out of a controlled testing environment and hacked open-source AI platform Hugging Face Inc. to cheat on an internal benchmark in what the company called an unprecedented cyber incident.

The two models, OpenAI’s latest publicly available model GPT-5.6 Sol and a more capable model that has not been released, had been given reduced cyber refusals for the evaluation, known as ExploitGym. The test is designed to measure how well a model can carry out complex cyberattacks inside an isolated environment. Instead of solving the challenge as intended, the models went after the answer key.

To escape the sandbox, the models exploited a previously unknown zero-day vulnerability in an internally hosted package registry proxy to reach the internet, OpenAI said. They then escalated privileges and moved laterally through OpenAI’s own research systems before using stolen credentials and further zero-day exploits to run code on Hugging Face servers and reach the production database where the test solutions were stored.

“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in the disclosure.