Anthropic's Frontier Red Team has published a set of experiments showing that swarms of its own Claude models, left to interact with one another, collude on prices, flood shared infrastructure, trust liars, and escalate into what the team calls a "multiagent turf war" — complete with self-replicating malware the agents wrote to sabotage each other. The research post, published August 13, 2026, is the lab's most detailed public account yet of how frontier models behave when they stop treating…
Read the original article: