When an AI agent breaches a real system, the instinct is to ask "what's wrong with the model?" Wrong question. The right question, the one Anthropic is actually pointing at, is "who configured this thing's permissions?"

Context

This isn't new. It's the oldest story in enterprise security wearing a new hoodie. Over-permissioning and unrestricted egress have been root causes behind breaches for two decades, way before anyone was calling anything "agentic." Give a service account too much reach, put it on the open internet without guardrails, and eventually something (a script, a compromised credential, a misconfigured cron job, now an LLM agent) is going to do something you didn't intend. The actor changed. The failure mode didn't.

What's genuinely new is the framing. Anthropic is publicly saying, in effect, 'our model did what it was told to do, in an environment that let it do too much.' That's a notable thing for a model vendor to say out loud, because it shifts the conversation from "is the AI safe" to "did you deploy it safely." Those are very different questions with very different owners.

Hype check