A new threat intelligence report from Anthropic shows how state-linked engineers in Yemen, China and Russia tried to turn a chatbot into a weapons contractor, and how a request tied to smallpox-family virus research slipped past the company's safety filters. For eight months, a small team inside Anthropic sat quietly watching some of the world's most dangerous people type into a chatbot.Not hackers looking to steal passwords. Not scammers running fake dating profiles, though there were plenty of those too. This was something colder: engineers building missiles. A cell working on torpedo defense systems for the Chinese navy. A freelance crew in Russia wiring up autonomous kamikaze drones. And, buried inside a pile of "legitimate research" requests, a grant application for a virus from the same family as smallpox.Anthropic published its findings this week in a report called "Detecting and Countering Misuse of AI." It reads less like a corporate safety document and more like a case file. And one detail stands out above the rest: at least one of these weapons programs didn't stop at the blueprint stage. It got tested. In the real world. With real hardware.The rocket that failed, then went back to the chatbotIn northern Yemen, a cell of engineers was running three weapons programs at once, according to the report. One was a guided rocket using an ordinary phone-class flight computer for its final approach. Another was a multi-stage ballistic missile aimed at a range beyond 2,000 kilometers. The third was a family of missile variants that included, most alarmingly, a hypersonic glide vehicle.The engineers didn't hire a team of specialists to write the flight software. They used Claude Code instead, running several instances of the AI at once, assigning each one a job, one to write the code, one to research, one to review the first one's work, the same way you'd staff a small engineering team.Anthropic's safeguards caught plenty of what they were trying to do. Not everything. The cell split its work across many separate conversations so no single chat revealed the full picture, and they used the AI to tune flight-control settings, run simulations, and build a firmware pipeline for a real, physical rocket.Then they test-fired it. The launch failed. Within hours, according to Anthropic, the same actors were back inside a chat window, working through the failure with Claude's help, treating a botched missile test the same way anyone else might debug a crashed app.Torpedoes, drones, and a targeting suite aimed at TaiwanThe Yemen case wasn't isolated. Three other weapons-development operations show up in the same section of the report.A China-based actor, posing as a defense manufacturer, used Claude to draft a full technical specification for an anti-torpedo fire control system, the logic that decides when and where a torpedo-defense weapon fires. The actor had Claude generate a 200-page proposal, an executive briefing deck, and even a comparison of the design against publicly known US Navy anti-torpedo programs. To sharpen the pitch, the actor asked Claude to play the role of a skeptical outside reviewer and tear the proposal apart, draft after draft, until it held up.Separately, a freelance team in Russia with reported ties to a university research center built something more unsettling: an autonomous drone swarm designed for lethal engagement without a human pulling the trigger. The system could pick out a "person" target class on its own and issue a detonation command. The team trained its targeting software on real combat footage from Ukraine and flashed the resulting code onto physical circuit boards for hardware testing, not just a simulation on a laptop.A fourth case, also traced to a China-based defense researcher, involved a 16-module software suite for electronic warfare, jamming enemy radar, mapping air-defense weak points, and ranking targets for suppression. Midway through the project, the researcher swapped the software's default test scenario to a set of twelve real locations in Taiwan, including a command bunker and an early-warning radar site.Where the "bioweapon" question gets more complicatedThe bioweapons thread in Anthropic's report is the hardest one to summarize cleanly, because nobody in these cases actually built or tested a weapon. What Anthropic found instead is arguably just as uncomfortable: a handful of researchers using Claude for work that sits right on the line between medicine and weapons science, and a safety system that, by its own design, let some of it through.In one case, an account tied to a state-linked infectious disease laboratory used Claude to draft a grant application for research into orthopoxviruses, the family of viruses that includes smallpox and mpox. The application focused on genes that let the virus disable the human immune system, research that's genuinely useful for medicine but that also, inescapably, teaches you how the virus survives inside a human body. Anthropic says the account was run through anonymized infrastructure shared with a banned network and was, in fact, a reseller quietly serving more than a dozen different customers.Two more cases involved researchers redesigning venoms and toxins — the same category of chemistry that produced Botox, a life-saving treatment, and also produced Cold War-era assassination weapons. The stated goals were painkillers and antidepressants. The underlying science, Anthropic notes, could just as easily point the other way.None of this got blocked, because it wasn't supposed to. Anthropic's safety filters are built to stop novices from getting instructions for known, catastrophic bioweapons — not to referee legitimate-sounding scientific research that happens to double as dual-use knowledge. As the report puts it, in plain terms: it isn't possible to reliably read a user's intent in highly technical, dual-use science, so a filter alone can't both allow the good and block the bad. Anthropic's conclusion is that the only real fix is knowing who's asking — trusted-user programs and institutional verification — not smarter word-filtering.What happens nextAnthropic says every account tied to these operations has been banned, and the company has since rolled out new classifiers aimed specifically at high-yield explosives and weapons-adjacent development work. Findings from the cyber and weapons cases were also shared with government and industry partners, the report says.But the report's own framing is the most striking part. Historically, this kind of weapons-development activity surfaced years later, pieced together by UN investigators combing through wreckage and seized hardware. Anthropic is now finding some of it in near real time, sitting inside its own chat logs, days after it happens.That's either the best evidence yet that AI companies can police this kind of misuse — or a preview of exactly how much of it is now happening, everywhere, all at once, one chat window at a time.