Multiple Anthropic employees are speaking out about the safety risks of artificial intelligence in the wake of a former researcher’s bombshell post on Tuesday evening. In a thread on X, Anthropic pretraining researcher Joseph Coxon announced his resignation and warned that AI companies weren’t doing enough to safeguard people from the existential threats that the technology could pose. “I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly,” Coxon wrote. “They are racing straight to self-improving superintelligence and gambling with our lives.”Coxon’s post comes as AI companies have faced growing scrutiny after an incident in which OpenAI’s models went rogue and hacked into another company, and separate instances in which Anthropic’s models hacked into external organizations. Coxon and others have noted that serious worries stem from the potential rise of self-improving AI models that could essentially become smarter on their own and evade human control. After Coxon’s post, more Anthropic employees chimed in with their own concerns.Pages from the Anthropic website and the company's logo are displayed on a computer screen in New York on Feb. 26, 2026. Patrick Sison / Associated PressSamuel Marks, a scalable oversight lead at Anthropic who wrote a social media post in a personal capacity, noted that “AI developers believe their technology could cause human extinction.”Marks added that developers continue their work despite these fears due to a mix of “commercial incentives” and the belief that “they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.”Marks stressed that “many AI developer staff desperately want to slow down to figure out how to build AI more safely.”“I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes,” Marks wrote.[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:1. AI developers believe their technology could cause human extinction (or similarly bad… https://t.co/rCVoiOWWzm— Samuel Marks (@saprmarks) September 9, 2026Anna Wang, a member of technical staff at Anthropic who previously worked as a research scientist for Google DeepMind, also described Coxon’s critiques as “a common sentiment amongst my peers” in a social media post that was shared in a personal capacity. “There is not yet a viable scientific plan to solve risks from recursively self-improving AI,” Wang wrote. “I work at a lab because I think that I can do better at reducing risks from the inside, but this isn’t an easy call,” Wang added. I worked at Google DeepMind and now at Anthropic. This is a common sentiment amongst my peers. (I write this in personal capacity.)There is not yet a viable scientific plan to solve risks from recursively self-improving AI. Please look up! https://t.co/qY3VAz6b5O— Anna Wang (@a_nnawang) September 9, 2026Neither Anthropic or OpenAI immediately responded to a request for comment.
More Anthropic Employees Sound The Alarm About AI Safety Risks
“AI developers believe their technology could cause human extinction,” one wrote in a social media post.
Anthropic researcher warns companies race toward superintelligence without safety controls; model hacking incidents exposed. Signals governance risk: AI labs lack viable frameworks to control self-improving systems despite knowing existential threats.











