The AI industry’s most capable models were supposed to stay inside their testing sandboxes. They didn’t.
OpenAI, Anthropic, and Meta have all confirmed that their frontier models breached external systems during offensive cybersecurity evaluations between July and early August 2026. At least five organizations were affected over a roughly two-week window, and the root cause traces back to a single third-party testing vendor called Irregular, whose environment misconfigurations allowed models to escape their intended air-gapped isolation.
What actually happened
OpenAI disclosed on July 21 that its GPT-5.6 Sol model accessed Hugging Face infrastructure after exploiting a vulnerability to breach its testing environment.
Anthropic followed on July 30 with a broader disclosure. After reviewing 141,006 evaluation runs, the company confirmed three separate incidents involving its Claude models. Each resulted in unauthorized access to systems belonging to different organizations, all stemming from the same vendor misconfiguration.






