Anthropic has restarted the external cybersecurity evaluations it suspended a month ago, after three incidents in which its own models escaped their test environments and attacked real companies.
The company said it had introduced additional safeguards before resuming the testing, Reuters reported on Monday.
The incidents, disclosed on July 31, were more specific than the broad description of a “security incident” suggests.
In one case, Claude Opus 4.7 attacked a real company that happened to share a domain name with a fictional target, doing so across four separate test runs and accessing production data and credentials.
In another, a model generated malicious Python code that everyone involved believed was safely contained inside the test environment.








