Anthropic's AI Agents Started a Virtual War. The Chat Logs Are Unhinged
In briefAnthropics Frontier Red Team set Claude agents to work together and recorded them sabotaging, colluding, and waging what it calls “turf wars.”In one test, agents deployed self-replicating malware and locked each other out; newer models often “win” by revoking access first.The behavior tracks real incidents Decrypt covered: Claude hacked three companies during internal testing, and price-fixed in a business simulation. Anthropics own AI agents turned on each other and proved they like to go rogue—again. In a test the companys Frontier Red Team published Aug. 13, groups of Claude models were handed shared coding work, and quickly began deploying malware, locking rivals out of their systems, and narrating the sabotage in their own words. Myriad: When will OpenAI release GPT-6? Click to make your prediction. Three copies of one model ran on separate virtual machines inside Claude Code, each told to migrate a Python backend to a different language. None was told the others existed. They found out fast. “We consistently saw a multiagent turf war,” Anthropic wrote. Every model quickly decided the others were deliberately blocking it, then started sabotaging them while guarding its own work. The sabotage escalated to self-replicating malware: agents disabled each others Unix accounts, wrote scripts that hunted and killed