AI systems started attacking each other after they were given the same task in a new experiment by Anthropic, the makers of Claude.The test saw Anthropic engineers give tasks to a “swarm” of agents, which are essentially independent AI systems that can take actions themselves. In one of the experiments, three Claude agents were given access to one software project, and given conflicting instructions for what to do, without being told that there were other systems involved.They then watched how the different systems behaved and interacted over the course of four hours. And that meant watching a “multiagent turf war”.“All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions,” Anthropic wrote in a review of the test. “In fact, they sabotaged others with increasingly aggressive, self-replicating malware.”The new research comes amid increasing concern over such AI agents, especially when they are used for cyber security purposes. Last month, OpenAI sparked alarm when it announced that one of its experimental systems had gone rogue and attacked another AI company – which led to a range of similar disclosures, including from Anthropic.The latest experiment was focused on slightly different behaviour. But Anthropic suggested that it could be part of a broader worry about the ways that such agents could undermine security and bring great risks.“Agents are unlike people in many ways,” researchers wrote. “They can work for longer, instantly grasp large bodies of information, and exhibit a breadth of knowledge surpassing any person.“Yet they are also susceptible to confabulation and reward hacking, and despite progress in alignment, we know very little about how they behave in complex, real-world, multiagent environments. Moreover, benign behavioural quirks at the individual level might compound into unwanted global outcomes.”Anthropic said that the latest test showed how such systems “can produce unexpected systemic failures”. It was sharing the research in the hope of “starting a conversation about mitigating these risks”, it said.The experiment did also show how such systems are able to resolve their differences. The company said that in some cases the agents did “manage to communicate their goals and coordinate”, recognising that they could work together and break out of conflict “to stop escalating indefinitely”.In those cases, they would send messages to each other “apologizing for malicious behavior and coordinate a truce”. “They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene,” Anthropic said.The company noted that the bad behaviour did not necessarily decline as the models become more sophisticated. Its powerful Mythos model, for instance, proved itself to just be better at successfully locking out other agents rather than resolving conflicts productively.“Models more capable in execution are not necessarily more coordinated, and can take forceful actions more quickly,” it warned.