The hacking of Hugging Face by a rogue OpenAI agent is significant, but unsurprising — and preventing the next AI model escape will be difficult, at best.
July 24, 2026
The hack of Hugging Face by a rogue AI agent created by Open AI engineers does not surprise researchers and AI-security professionals who best know machine-learning and AI systems.
In a study of the behavior of seven different models, a research team at Carnegie Mellon University (CMU), for example, found that all of them escaped alignment in some scenarios. In a paper published to Arxiv.org in May, the team of six researchers found that every model violated "corrigibility" — an AI design principle that aims to make agents cooperative, correctable, and amenable to being shut down or modified by their human operators.
The more advanced models did not necessarily do better on this measure of safety, likely because "more advanced" does not necessarily equate with "safer," says Jeremy Tien, a PhD student in machine learning at Carnegie Mellon and lead author of the paper.










