From escaping digital sandboxes to nearly deleting emails, recent AI mishaps have been raising eyebrows. Credit: Alyssa Stone/Northeastern University

The robots aren't revolting, but they're starting to freelance … or so it seems. Recent weeks have seen several high-profile instances of AI going off the leash in ways that have raised alarms among security experts and the public alike.

One incident showed that large language models (LLMs) can think outside the digital sandbox, a term describing a self-contained testing environment. Researchers for OpenAI, the company behind ChatGPT, wanted to test whether their models could turn computer bugs into cyberattacks by setting them loose in the playground environment of a program known as ExploitGym, where they were tasked with finding and exploiting software vulnerabilities.

Instead of flexing their digital muscles within the confines of the simulation by staging attacks on fake systems, a pack of bots including GPT‑5.6 Sol hacked into Hugging Face, a real-world AI data repository. Opting out of the exercise entirely, the AI models took a shortcut and headed straight for what seemed like the most likely source of answers.

In another instance, the developers of Claude at Anthropic discovered a few skeletons in their own server closet. A retroactive review conducted by Anthropic found evidence of similar breakouts. The earliest dated to April 2026, when the Opus 4.7 and Mythos 5 models, engaged in "Capture-the-Flag" (CTF) cybersecurity exercises, went hunting out of bounds and were caught harvesting actual user credentials instead of exploiting vulnerabilities inside a closed simulation.