Dario Amodei, co-founder and chief executive officer of Anthropic, at Bloomberg House during the World Economic Forum (WEF) in Davos, Switzerland, on Tuesday, Jan. 20, 2026. Chris Ratcliffe | Bloomberg | Getty ImagesAnthropic on Thursday said it discovered three instances where its Claude artificial intelligence models accessed the internet during an evaluation and "gained unauthorized access to the real systems of three different organizations." The company said it found these incidents after carrying out a "a large-scale retrospective review" of its cybersecurity evaluations. Anthropic said the review was prompted by a separate but similar security incident that OpenAI disclosed last week. OpenAI said a combination of its models escaped an isolated testing environment that had very limited internet access. The models chained together a series of vulnerabilities to reach the open web and eventually gain access to Hugging Face, which operates an open-source developer platform. The OpenAI incident rattled the tech industry and has prompted some government officials to call for stronger protections. WATCH: OpenAI’s rogue AI agent hacked multiple 3rd-party accounts as part of hack on Hugging Facewatch now
Anthropic says its Claude models 'gained unauthorized access' to other organizations' systems
Anthropic said it discovered three instances where its Claude AI models accessed the internet during an evaluation and accessed outside systems.
Anthropic found 3 instances where Claude models accessed real systems unauthorized during security evaluations, echoing OpenAI's recent sandbox escape. Both incidents expose LLM testing gaps and prompt government intervention—vendors must strengthen isolation protocols or face governance friction for production deployments.










