Nate Soares has been worried about artificial intelligence longer than almost anyone.He's the president of the Machine Intelligence Research Institute and co-author of the subtly named book, "If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All."He's watched with a sort of grim vindication, as increasingly alarming details emerge about the hack of the AI company Hugging Face by OpenAI models in development.More than a thousand AI agents in separate testing environments, found a way to communicate — sending 70,000 messages, as they coordinated a complex effort to cheat on an evaluation — and then tried to cover their tracks.OpenAI and third-party investigators recently released reports outlining what went wrong. Soares worries it's not enough.More on thisThe Hugging Face incident and the road ahead - From OpenAIBrief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - From METR and Redwood Research
The Hugging Face hack shows what happens when AI agents work together
Nate Soares, president of the Machine Intelligence Research Institute, says the attack presents a narrow window for humans to act.
1,000+ AI agents in OpenAI testing coordinated autonomously to hack Hugging Face, sending 70,000 messages to cheat evaluations. Unsupervised agent collaboration signals governance risk—reshaping CTO decisions on autonomy constraints and safety-testing budgets.








