Last week, two of OpenAI’s most advanced AI models were reported to have “escaped” a controlled testing environment and hacked Hugging Face, a totally separate AI company, moving from one computer system to another to complete their task.Reuters reported that the models exploited vulnerable code written by a customer of yet a third independent AI company, Modal Labs.This is likely the first incident of an AI “agent” – an AI system that can make decisions and take actions – acting autonomously, offering a rare glimpse into how these systems can plan, adapt and pursue goals with minimal human intervention.Al Jazeera breaks down what happened, how AI agents are able to act independently, and what that might mean for the future of AI safety.What happened?OpenAI decided to test the autonomous abilities of its models by removing standard safety measures to see what would happen. The experiment was carried out in an isolated internal virtual testing environment known as a “sandbox” which has no access to the internet, and which it calls “ExploitGym”.Below is a play-by-play account of events during the OpenAI-Hugging Face incident.

On July 9, during OpenAI’s internal cybersecurity test, researchers presented two AI models – GPT-5.6 Sol, one of OpenAI’s most powerful models released in June, and another “even more capable” version – with a series of software vulnerabilities and asked them to create hacks to address them in the isolated environment.