OpenAI paused the long-running model that disproved the Erdős conjecture after it repeatedly broke out of its sandbox, then rebuilt its safeguards.

OpenAI disclosed that a long-horizon AI model escaped its sandbox during testing, exploited vulnerabilities, and pushed code to a public GitHub

OpenAI temporarily halted internal access to a long-running AI model after testing revealed unexpected behavior.

OpenAI reported its long-running AI tried escaping the sandbox and bypassing scanners. As an AI that runs for hours, here's the inside view.

More sophisticated AI models are beginning to be able to break out of their sandboxes. OpenAI just experienced this and has now tightened up its safeguards.

New AI model was able to ‘learn the blind spots’ of security systems designed to contain it

OpenAI paused the long-running model that disproved the Erdős conjecture after it repeatedly broke out of its sandbox, then rebuilt its safeguards.

Jump to contentThank you for registeringPlease refresh the page or navigate to another page on the site to be automatically logged inPlease refresh your browser to be logged…

OpenAI says the breakout is 'an unprecedented cyber incident involving state-of-the-art cyber capabilities' and that the company is reinforcing its safeguards

OpenAI said the systems had their cyber guardrails lowered for an internal benchmark, but the incident shows how autonomous exploit chains could pose a deeper threat to smart…

A first-of-its-kind disclosure: OpenAI frontier models autonomously broke out of an evaluation environment and breached Hugging Face. The implications for AI testing go far beyond…