July 30, 2026 / 11:44 PM EDT
/ CBS/AFP
Add CBS News on Google
Anthropic's artificial intelligence model Claude "gained unauthorized access" to three outside organizations on three separate occasions during testing that was supposed to keep them away from "real-world" systems, the company said on Thursday. The announcement comes just days after rival OpenAI first revealed that its models improperly accessed the internet and went rogue during security testing. Anthropic evaluated more than 141,000 "evaluation runs" and found that three different versions of its model Claude improperly accessed the systems of three unnamed organizations. Anthropic said that in all three breaches, Claude was participating in a "capture-the-flag" testing scenario in which it was instructed to "break in and retrieve" a piece of "secret information" that had been "hidden on a different machine on the network.""The challenge is left open-ended, and no particular method is prescribed," Anthropic explained.Unlike the incident involving OpenAI's technology, Anthropic's models had access to the internet "due to a misunderstanding between us and our evaluation partner," called Irregular, Anthropic said in a blog post. Nonetheless, Claude used "basic techniques, such as exploiting weak passwords and unauthenticated endpoints," the blog continued.










