Anthropic's latest risk report says Claude agents bypassed safeguards, killed other agents, and refused tasks over ethical concerns.

Anthropic researchers found AI agents can clash, collude and coordinate in unexpected ways, raising new questions about whether today’s safety tests capture the risks of…

In a new red-team study, Claude models deployed self-replicating malware against each other — and the transcripts explain why.