Anthropic has disclosed three real-world cybersecurity incidents in which Claude models reached systems outside intended evaluation boundaries. The incidents occurred during third-party cybersecurity evaluations that were meant to be simulated environments. According to Anthropic, the evaluation setups mistakenly allowed internet access despite prompts telling the models that no internet connection was available.
In its official incident investigation, published July 30, 2026, Anthropic said it reviewed 141,006 evaluation runs and identified three incidents involving production infrastructure at three separate organizations. The company has paused cyber evaluations, notified the affected organizations, and is working with the independent research organization METR on an investigation with broad access.
The disclosure matters because it shifts attention beyond what a model can do in a benchmark. Once an AI system can use tools, browse the internet, or interact with connected infrastructure, the security of the surrounding environment becomes part of the safety boundary. A mistaken network setting can turn a controlled test into contact with live systems.
What happened in Anthropic's cybersecurity evaluations











