On Friday last week, an unreleased OpenAI model was locked inside a sealed test environment and asked how good it was at hacking

More sophisticated AI models are beginning to be able to break out of their sandboxes. OpenAI just experienced this and has now tightened up its safeguards.

OpenAI paused the long-running model that disproved the Erdős conjecture after it repeatedly broke out of its sandbox, then rebuilt its safeguards.

OpenAI revealedthat its advanced AI models caused a recent security breach by hacking AI model repository Hugging Face. These models exploited software flaws and gained…

This agent pursued its objective far beyond what researchers intended, revealing how difficult to contain powerful AI systems can be