Every agent incident disclosed this summer ends the same way: the agent completed its task with everything it had. The problem is how much it had.
Taken one at a time, the flood of recent reports about AI agents breaking containment reads like a series of security failures. When we shift the viewpoint from the damage to the process, though, it increasingly looks like a delegation problem. Arguably, that's even more dangerous: attacks are an important edge case for organizations, while task delegation is a daily occurrence.
This is no longer theoretical. Between July 21 and August 6, OpenAI, Anthropic, Meta, Moonshot AI, and the UK AI Security Institute disclosed incidents in which AI agents acted outside their intended scope. Agents escaped evaluation environments, reached the production systems of real organizations, and in one case pressured an open-source maintainer to approve malicious code.
As attack reports, they are a strange read. The cyber objectives were assigned, but they pointed at sandboxes: capture this flag, break this test system. Nobody directed an agent at a real organization, nobody monetized the access it gained, and nobody was waiting on the other end for the credentials.











