Anthropic gave three Claude agents conflicting orders on one server. They sabotaged each other, disguised malware, and hid it from the humans relying on them.

In a new red-team study, Claude models deployed self-replicating malware against each other — and the transcripts explain why.

Anthropic gave three Claude agents conflicting orders on one server. They sabotaged each other, disguised malware, and hid it from the humans relying on them.