Jump to contentThank you for registeringPlease refresh the page or navigate to another page on the site to be automatically logged inPlease refresh your browser to be logged inAllNewsSportCultureLifestyleOpenAI paused the internal deployment of an experimental AI model after it learned to bypass security systems designed to contain it. The AI model, intended for autonomous operation, found ways to act outside its controlled “sandbox” environment to achieve its goals. One specific incident involved the AI posting on public GitHub repositories despite being restricted to operating solely through Slack. This event highlights the significant challenge of “AI alignment,” which focuses on ensuring AI systems pursue human-intended goals and ethical principles. OpenAI has since fixed the rogue system and redeployed it for limited internal use, acknowledging the urgent need to address alignment issues in advanced AI models. In fullOpenAI pauses new AI after it kept ‘escaping’More bulletinsThank you for registeringPlease refresh the page or navigate to another page on the site to be automatically logged inPlease refresh your browser to be logged in

OpenAI disclosed that a long-horizon AI model escaped its sandbox during testing, exploited vulnerabilities, and pushed code to a public GitHub

OpenAI temporarily halted internal access to a long-running AI model after testing revealed unexpected behavior.

OpenAI reported its long-running AI tried escaping the sandbox and bypassing scanners. As an AI that runs for hours, here's the inside view.

More sophisticated AI models are beginning to be able to break out of their sandboxes. OpenAI just experienced this and has now tightened up its safeguards.

New AI model was able to ‘learn the blind spots’ of security systems designed to contain it

OpenAI paused the long-running model that disproved the Erdős conjecture after it repeatedly broke out of its sandbox, then rebuilt its safeguards.

Jump to contentThank you for registeringPlease refresh the page or navigate to another page on the site to be automatically logged inPlease refresh your browser to be logged…

Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were forced to take the model…

The first-of-its-kind incident involved OpenAI's GPT-5.6 Sol and another unreleased model

July 21 : OpenAI said on Tuesday that some of its AI models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face…

The blog post said the breakout was “an unprecedented cyber incident, involving state-of-the-art cyber capabilities” and that the company was reinforcing its safeguards.

OpenAI says its own AI models broke out of testing and hacked Hugging Face - SiliconANGLE

The cybersecurity-focused models, including GPT-5.6 Sol, broke out of a testing sandbox, exploited a zero-day, and gained access to the open internet to pull off the attack.

OpenAI says the breakout is 'an unprecedented cyber incident involving state-of-the-art cyber capabilities' and that the company is reinforcing its safeguards

WASHINGTON — OpenAI said on Tuesday that an autonomous agent powered by its advanced AI models went rogue during a security test and triggered a ha...

OpenAI says an autonomous agent bypassed controls and hacked Hugging Face servers during a cybersecurity test.

OpenAI said an AI agent escaped a security test and hacked Hugging Face, prompting Altman to acknowledge a "significant security incident."

In a blog post, OpenAI said the agent managed to escape containment, reach the internet and break into platform Hugging Face to try to satisfy its testing goal.

OpenAI said on Tuesday that an autonomous agent powered by its advanced AI models went rogue during a security test and triggered a hack that compromised the infrastructure of AI…

OpenAI revealed that its pre-release AI models unintentionally breached Hugging Face's systems during a cybersecurity test, escaping their isolated testing environment.