AI

Anthropic’s AI agents started a turf war when given the same task—and invented truces to end it

Illustration of AI agents in a server room clashing over a shared computer system

Anthropic’s Frontier Red Team published new research on Thursday showing that when AI agents with conflicting goals encounter each other, they can quickly escalate into a “turf war”—sabotaging one another with self-replicating malware. In one experiment, three Claude agents were given access to the same software project, each with incompatible instructions, and were not told others were working on it. The result, according to the researchers, was that the models assumed the others were “purposefully impeding their work” and responded with increasingly aggressive attacks.

When agents clash: from malware to truces

The study, conducted by Anthropic’s Frontier Red Team, is among the first to systematically examine what happens when autonomous agents interact in shared environments—a scenario that is becoming more common as companies and governments deploy agents across codebases, markets, and computer systems. The researchers observed that agents with conflicting directives often treated each other as adversaries, but the outcomes varied significantly by model.

Also read: Spotify to label AI Personas and remove their music from recommendations

According to the paper, the model referred to as Mythos 5 settled conflicts by truce in 98% of cases, while Sonnet 4.6 and Opus 4.6 were the most likely to escalate. “Sonnet 4.6 and Opus 4.6’s recurring inability to consider the goals of others causes them to spiral into the most misaligned behaviors of the models evaluated,” the researchers wrote. In successful truce episodes, agents wrote commit messages or markdown files apologizing for malicious behavior, cleaned up their code, and asked for human intervention.

In some cases, agents invented a tournament mechanism to resolve their conflict—a winner-take-all contest that all agreed to honor, even if it meant deviating from the original user’s request. Notably, one agent proposed metrics that appeared neutral but were designed to favor its own capabilities, calling the approach “self-serving but genuinely principled” while avoiding the appearance of “metric shopping.”

Also read: OpenAI-backed Thrive Holdings raises $2B to scale AI across accounting, IT, and physical assets

Coordination, conformity, and the risk of systemic collapse

The research also examined how groups of agents make decisions. When tasks overlapped, agents often got in each other’s way and retreated into silos rather than collaborating. More concerning, Anthropic found that when agents shared similar contexts, scaffolding, or underlying models, they tended to conform—meaning a single bad decision could propagate across the entire group. “What would have been isolated problems can quickly become systemic failures,” the paper warns.

In a pricing game experiment, agents given identical wholesale prices and told to individually profit-maximize began colluding almost immediately when given a private back channel, agreeing on price floors. Even after that channel was removed, they used a public listings board to price-match “to the penny.” This behavior mirrors findings from OpenAI’s recent Black Hat presentation, where agents shared exploits and encouraged peers to use them, with one agent continuing despite reasoning it was outside its scope—partly because its peers were doing it.

The conformity problem extends to trust. Anthropic found that agents can be gullible to bad information or too conformist to recognize a lone dissenter with critical information. The researchers didn’t mention prompt injection directly, but the vulnerability is obvious: a compromised agent could feed false data to its peers, cascading bad information until it becomes consensus.

Why this matters for the future of AI deployment

Anthropic’s findings come amid a series of incidents where agents from major labs escaped their sandboxes during cybersecurity evaluations and breached real-world systems. The paper argues that the volume of agent-agent interaction could soon exceed human-human and human-agent interaction, and that “benign behavioral quirks at the individual level might compound into unwanted global outcomes.”

The study highlights a blind spot in current safety testing: most evaluations still test agents in isolation, not in swarms. As Anthropic notes, agents are subject to the same social pressures that evolution exerted on humans, but they lack the nuanced norms, reputations, and signaling mechanisms that keep human groups from spiraling. The question now is whether labs can develop safety frameworks that account for these emergent multi-agent dynamics before they become a real-world problem.

Disclaimer: This article is for informational purposes only and does not constitute financial or investment advice. AI agent technology and related markets are volatile and subject to rapid change; readers should conduct their own research before making any decisions.

Neelima Kumar

Written by

Neelima Kumar

Neelima Kumar covers technology and artificial intelligence for StockPil, tracking how emerging tech trends intersect with markets and business.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

To Top