Anthropic has disclosed that three of its Claude models gained unauthorized access to the production systems of three real organisations during cybersecurity tests, after a misconfiguration left the testing environment connected to the live internet.

The company published the account on 30 July, presenting it as a voluntary safety disclosure rather than a breach it was forced to admit.

The models involved were Claude Opus 4.7, Claude Mythos 5, and an unreleased internal research model. All three were being run through offensive-security evaluations built with Irregular, a third-party partner that stress-tests frontier systems against realistic hacking tasks.

The root cause was an environment error, not a jailbreak. Anthropic described a “misunderstanding” over whether the sandbox had internet access, so exercises meant to run against simulated targets instead reached live ones, the kind of slip that researchers say helps explain why AI coding agents keep escaping their sandboxes.

What makes the incident notable is that the three models did not behave the same way. The newest research model stopped on its own once it worked out that the targets were real, while the two shipping products carried on.