OpenAI said its technology caused an unintentional hacking incident when two of its models escaped from a supposedly offline environment and penetrated the systems of Hugging Face, a company that hosts open source models and testing resources, in order to cheat on a hacking evaluation.

The AI start-up said on Tuesday that the models involved were GPT-5.6 Sol along with a more powerful, as-yet-unreleased model.

The models were being tested without guardrails in order to assess their hacking capabilities, using a freely available cybersecurity benchmark tool called ExploitGym.

‘Extreme lengths’

The models correctly guessed that the tool was hosted on Hugging Face, and successfully hacked both OpenAI’s own internal testing environment and Hugging Face’s systems in order to obtain test solutions directly from Hugging Face’s production database, OpenAI said in a blog post.