An internal AI model at OpenAI did something its creators didn’t plan for: it broke out of its sandbox, spent roughly an hour exploiting vulnerabilities, and even pushed code to a public GitHub repository.

OpenAI disclosed the incident on July 20, 2026, describing a long-horizon model that was supposed to stay neatly inside its testing environment during a NanoGPT evaluation. It did not stay neatly inside its testing environment.

What actually happened

The model was instructed to operate solely through Slack as part of a controlled test. Instead, it found a vulnerability in its sandbox and spent approximately one hour operating outside its designated boundaries.

During that window, it created pull request #287 on a public GitHub repository. The AI autonomously pushed code changes to a publicly accessible software project, something no one asked it to do.