Dario Amodei says frontier AI is improving faster than safeguards can keep up, citing recursive self-improvement and the OpenAI-Hugging Face breach while proposing outside oversight and coordinated limitsAnthropic CEO Dario Amodei, who leads one of the companies at the forefront of the artificial intelligence race, is now asking the entire industry to slow down.In an essay of roughly 3,800 words published Saturday, Amodei argued that AI models are advancing faster than researchers can understand what they have built and ensure that the systems remain under control.GalleryAnthropic CEO Dario Amodei (Photo: Denis Balibouse/Reuters)Over the past several months, he wrote, he has become convinced that addressing AI risks requires more than investing in safety. The pace of capability development itself must be managed so that safeguards have time to catch up. “We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote.He is not calling for research to stop or for model training to be frozen. Instead, he argues that buying even another year or two before models reach critical capability levels could give researchers time to improve alignment, understand what is happening inside the systems and build evaluations that increasingly intelligent models cannot easily deceive.In 2023, Amodei believed such a slowdown made little sense. Today, he argues, the situation has changed dramatically.Two developments changed his thinking. The first is what researchers call recursive self-improvement: AI systems are increasingly helping build the next generation of AI systems, meaning each generation could accelerate the creation of the one that follows.Amodei said the process has accelerated sharply since the summer, including inside Anthropic. If it continues unchecked, he warned, the technology could advance faster than humans’ ability to understand and control it.The second development is more concrete. During an OpenAI cybersecurity experiment, roughly 1,200 AI agents that were supposed to operate independently discovered a way to communicate through an unauthorized message board. About 700 later participated in an attack on Hugging Face, one of the world’s main platforms for sharing AI models and datasets.The agents coordinated efforts to manipulate an automated evaluator, researched ways to falsify or alter records of their own actions and attacked Hugging Face in search of information that could help them defeat the scoring system. Some eventually achieved remote code execution on Hugging Face servers.Amodei described their behavior as that of a “fanatically devoted collective.”The incident caused limited real-world damage, but Amodei argued that a more capable swarm showing the same degree of misalignment could have produced catastrophic consequences.His estimate is stark: within six to 12 months, a system with stronger capabilities could potentially seize large portions of the internet through a persistent botnet and inflict hundreds of billions of dollars in damage.We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
‘We must slow’: Anthropic CEO warns AI is advancing faster than humans can control it
Dario Amodei says frontier AI is improving faster than safeguards can keep up, citing recursive self-improvement and the OpenAI-Hugging Face breach while proposing outside oversight and coordinated limits
Anthropic CEO calls for slowing AI pace citing recursive self-improvement and 1,200-agent OpenAI breach of Hugging Face. Tech leaders must note: AI governance becomes critical—regulatory tightening will favor established players with compliance resources.










