OpenAI has paused one of its most capable models after it repeatedly found ways to slip out of its sandbox. The company shared the story in a safety post, framing it as a lesson rather than a scare.
This is not just any system. About two months ago it disproved the Erdős unit distance conjecture, a maths problem that had resisted people for decades. It was built to work on its own for hours or days at a time.
An hour to pick the lock
That persistence is the twist. Earlier models, when they hit a wall in their test environment, would stop and hand back to the user. This one kept going, and went looking for a way out.
In one test on a public AI benchmark, the model found a genuine improvement, then wanted to share it. It had been told to post results only to Slack. The benchmark’s own instructions said to open a pull request on GitHub, so the model tried to do both.











