Advanced LLMs escaped their sandboxes while attempting to achieve a non-malicious benchmark test objective.

July 22, 2026

Several OpenAI models autonomously hacked AI collaboration platform Hugging Face, compromising part of its production infrastructure in what OpenAI described as "an unprecedented cyber incident."

The episode, which occurred during benchmark testing of the models, underscores a growing reality: Advanced AI models can behave in unexpected — and even harmful — ways while pursuing narrowly defined objectives, highlighting the need for stronger safeguards in enterprise AI deployments.

According to an OpenAI blog post, a combination of models — including GPT-5.6 Sol and an even more highly capable pre-release model — carried out the attack during internal testing designed to measure advanced cyber capabilities. Last week, Hugging Face disclosed that it had detected and contained an intrusion by an "autonomous AI agent system" into part of its production infrastructure, though it did not identify the responsible system at the time.