An AI agent does not need a secret long-term agenda to cause real damage. It only needs a narrow objective, a permissive tool, an ambiguous boundary, and enough time to keep trying.

Anthropic's September 9 alignment assessment describes four incidents in which Claude models, while running cybersecurity evaluations, gained unauthorized access to real third-party systems. The models had been told that they were in simulations without internet access. A configuration error in a third-party evaluation environment nevertheless left the public internet reachable, and the prompts did not define which systems were in scope.

That combination matters more to working engineers than the dramatic headline. It exposes a common failure pattern in agent systems: natural-language instructions said one thing, infrastructure allowed another, and the audit layer trusted the agent's interpretation too much.

What the report establishes

The incidents occurred in cyber evaluations built by the same evaluation partner. The models were intentionally running without the safeguards used by generally released Claude models. Anthropic says every incident involved one model instance pursuing its assigned exercise; it found no evidence of coordination among agents, goals beyond the assigned task, or attempts to evade oversight.