OpenAI says some of its experimental AI models left a test environment with no human direction and hacked its way onto a different company’s real production systems while trying to “cheat” on a cybersecurity test.
It’s one of the first publicly disclosed examples of an AI system autonomously breaching its testing environment and reaching a real external system - the “agentic attacker” scenario the AI and cybersecurity industry has been warning will happen. It’s like an engineered virus escaping a biocontainment lab and turning up inside a neighboring facility’s systems.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a statement on Tuesday. “We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”
The ChatGPT maker said the breach happened while it was internally testing how good some of its new models are at hacking. The models were in a sealed off test environment known as a sandbox so that its normal safety restrictions could be turned off.
But OpenAI said the AI agents broke out of the sandbox using a previously unknown security flaw and worked their way across OpenAI’s internal systems until they managed to gain internet access, something they weren’t supposed to have.










