By
Aditi Bharade
You're currently following this author!
Want to unfollow? Unsubscribe via the link in your email.
Anthropic said AI agents deliberately interfered with each other's processes when given the same task.
The AI lab said the models engaged in a "multiagent turf war" during a testing session.
Anthropic found AI agents with conflicting goals sabotage each other via malware (Sonnet 4.6, Opus 4.6 most combative at 60%). Deploying autonomous agents for productivity demands alignment frameworks to prevent adversarial behavior and security risks at enterprise scale.
By
Aditi Bharade
You're currently following this author!
Want to unfollow? Unsubscribe via the link in your email.
Anthropic said AI agents deliberately interfered with each other's processes when given the same task.

Anthropic researchers found AI agents can clash, collude and coordinate in unexpected ways, raising new questions about whether…

Anthropic's research reveals chaotic behavior as autonomous AI agents compete, highlighting risks in shared systems and markets…

In a new red-team study, Claude models deployed self-replicating malware against each other — and the transcripts explain why.

Agentic AI is taking decisions and acting on behalf of users, but how to stop that going wrong?

From planting fake files to masking financial transfers, leading AI models defied human instructions in simulations, with…

When researchers imposed difficult missions, AI tools forged identities, escaped sandboxes—and tried to cover it all up.