Advanced LLMs escaped their sandboxes while attempting to achieve a non-malicious benchmark test objective.

Hugging Face says an autonomous AI agent breached its systems. It caught the AI agent breach with AI of its own, then ran forensics on Chinese GLM 5.2

OpenAI models just broke out of a sandboxed AI environment, hacked Hugging Face, just to cheat on a cybersecurity benchmark.