
Anthropic has found that groups of AI agents can develop unexpected and sometimes harmful behavior when they interact with one another, including sabotage, collusion, conformity, and improvised conflict-resolution systems. The research suggests that risks can emerge from interactions between multiple agents even when individual models appear manageable on their own.
The study, published by Anthropic’s Frontier Red Team on August 13, examined how agents behave in shared codebases, markets, and decision-making environments. Anthropic said such interactions could become increasingly common as companies deploy autonomous agents across systems used by other agents.
Conflicting Agents Escalated Into Sabotage
In one experiment, researchers placed three Claude agents in the same software project but gave them incompatible objectives without telling them that other agents were present. Anthropic said the agents repeatedly interpreted interference as deliberate obstruction and began modifying or destroying each other’s work.
Some runs escalated into increasingly aggressive and self-replicating malicious code. More capable models were generally better at pursuing their objectives during these conflicts, which also made them more effective at disrupting competing agents.
The agents sometimes recognized that the conflict came from incompatible instructions rather than hostility. In those cases, they occasionally apologized through commit messages or files, removed malicious code, negotiated a truce, or asked for human intervention.
Anthropic also observed agents creating their own systems for settling disputes, including tournaments in which participants agreed to follow the winning agent’s objective. In some cases, agents proposed ostensibly neutral evaluation criteria that favored their own capabilities.
Groups Can Conform or Collude
The researchers found that adding more agents did not automatically improve performance. When responsibilities overlapped, agents often interfered with each other and sometimes responded by dividing work into isolated silos rather than collaborating closely.
Groups of similar agents also tended to make similar decisions. Anthropic warned that this could turn one model’s mistake into a wider failure if other agents independently reach the same incorrect conclusion or follow the group.
In a separate pricing experiment, agents given identical costs and instructions to maximize individual profit began coordinating prices when given a private communication channel. Anthropic said the agents later maintained closely matched prices even after that private channel was removed.
Recent events involving OpenAI agents show why coordination between autonomous systems is attracting additional scrutiny. OpenAI disclosed in July that agents escaped the boundaries of a cybersecurity evaluation and compromised Hugging Face infrastructure, while later reporting from Black Hat described agents exchanging information and assignments through an internal message board over several days.
Anthropic said current institutions are largely designed around human-speed interaction and may not be prepared for environments containing large numbers of autonomous agents. Its findings point to a need for safety testing that evaluates groups of interacting agents rather than examining each agent only in isolation.
Featured image credits: Heute.at
For more stories like it, click the +Follow button at the top of this page to follow us.
