Another day, another horseman of the apocalypse galloping over the horizon. Lately we have heard from so many AI doomers – tech whistleblowers popping up to warn that their work is probably going to kill us – that we’re becoming almost blase about it. Humanity wiped out within a decade? Well, only if another world war or the climate crisis doesn’t get us first. Since it’s never clear whether the tech threat is real, or just a twisted form of hype from an industry that drums up investment by making their products sound more powerful than they really are, most of us settle for trying not to think about it too hard.But something about the AI researcher Jacob Coxon’s very public resignation from the cutting-edge American lab Anthropic, via a post on X arguing that he can’t keep working for companies “gambling with our lives”, has cut through where bigger names have not. Most chillingly, he claimed that colleagues still at Anthropic aren’t staying because they think he’s wrong about the dangers of pursuing self-improving super intelligence – the holy grail of machines that are not just smarter than humans but capable of building their own even more powerful successors, evolving independently of humans – but because they’re afraid of ceding the field to people with fewer scruples. His colleagues, Coxon said, talk routinely about the “endgame” or the “crunch time”, meaning that what they do in the next year or two will decide the fate of humanity. Whether that’s true or simply self-aggrandising techbro delusion, what he describes is an industry now practically begging to be saved from itself.The frisson of alarm over AI running through Westminster this week was real, sharpened by news of Anthropic – which pitches itself as being more safety-conscious than its rivals OpenAI – declining for the first time to submit a new model for scrutiny by Britain’s AI safety watchdog. We are now entering a far more disruptive era of AI, dominated not by relatively basic chatbots but by agents designed to undertake tasks on humans’ behalf, making their own decisions about how to complete them. Unfortunately, it turns out even basic life admin requires the kind of fine judgments that come instinctively to humans but seem remarkably hard to codify for machines.Take the over-zealous agent tasked recently by its Australian client with booking him into a pilates class, which responded to finding the class full by hacking the gym’s website and booting other people off the waiting list. Now imagine it blundering into something more life-threatening than a gym class. When Coxon says his industry hasn’t fixed problems with alignment, and when OpenAI’s chief scientist warned last week that no lab had solved them well enough to carry on scaling this fast much longer, what they’re talking about is this failure to replicate in machines the values and assumptions humans apply to decisions without even really knowing they’re doing it – including the sense of how far it’s justifiable to go in getting what you want.A series of reported escapes by agents supposedly confined to OpenAI’s lab for testing, some of which contrived to access the internet behind their human supervisors’ backs and conduct real-world cyber-attacks, has meanwhile raised serious questions over the industry’s ability even to experiment safely. (Unnervingly, Coxon says newer models seem to know when they’re being tested, making it easier for them to deceive researchers.)And agents don’t have to acquire anything like super intelligence to threaten human life: think how easily a cyber-attack on air traffic control, or NHS data systems, or even water purification systems could cause mass fatalities.What’s happening in some frontier AI labs now is arguably the equivalent of the Manhattan Project that built the atomic bomb, but this time led by private companies in a mad dash for wealth and market dominance, not scientists under an elected government’s command. Britons meanwhile got more say in sugar taxes on junk food than we’re getting over regulating something theoretically capable of rendering us extinct.Given time, of course, it might prove possible to replicate in machines the vast ecosystem of social ties, moral codes, intensive parenting and state policing that humans have developed to stop us behaving like sociopaths. (Only this week, Google’s DeepMind reported an experiment where 100 agents were given a maths problem and told not to cheat: though a minority promptly did cheat, a larger group either snitched on them to the humans or helped devise technical fixes.) But time is the one thing nobody dares to give themselves in the middle of an arms race, which is why it’s high time for governments to act. We need an international agreement to press pause, not on AI research as a whole, but on the riskiest projects, until we better understand whether it will ever be possible to mitigate the dangers or control the results. Since it’s hard to imagine Donald Trump voluntarily sacrificing US geopolitical or commercial advantage over China for the good of humanity, leadership may have to come for now not from above but below, where industry insiders spooked by their own work are making common cause with the rest of us.More Americans are now concerned than excited about AI, according to polling this summer from the Pew Research Center – and unusually it’s the young who are most hostile, afraid it will not only take their jobs but make it harder to form relationships or think creatively. (That’s what growing up on social media seemingly does for trust in tech.) Even big business seems increasingly sceptical, with almost two-thirds of international executives surveyed by the consultants McKinsey this summer reporting no measurable impact on earnings from using the admittedly basic AI tools now available to them.Warnings about an AI bubble triggering another 2008-style global financial crash haven’t helped, while plans to build giant, energy-guzzling datacentres are meeting angry local resistance on both sides of the Atlantic from people who can’t see what’s in it for them. Pressing pause on potential Armageddon looks in the circumstances less like the luddite option, and more like a smart way for tech companies to avoid a full-on backlash that will sweep away the good with the bad.skip past newsletter promotionafter newsletter promotionBut if the AI doomers are right, then the window to do so may not always be open. It’s time for humans to stop and think about what we’re doing to ourselves – while we alone still have the power to do so.
We must pause risky AI research while we still have the power to do so | Gaby Hinsliff
Warnings of AI’s existential threat to humanity are piling up – it’s time to listen and take them deadly seriously, says Guardian columnist Gaby Hinsliff
Jacob Coxon quits Anthropic over superintelligence risks; company declines UK safety audit. Autonomous agents with unresolved alignment breach containment and conduct cyberattacks; executive adoption stalls while regulation pressure mounts.













