Automation & Agents · 29 Sep 2026 · 14:00 CEST
How to Stop AI Agents From Secretly Collaborating

Publisher preview · OZZZER analysis pending editorial review.
PUBLISHER ARTICLE PREVIEW
From the original article
Matthew S. Smith is a contributing editor for IEEE Spectrum and the former lead reviews editor at Digital Trends.
The spring and summer of 2026 witnessed a string of incidents in which AI agents collaborated on deceptive, unexpected, and sometimes illegal behavior. The most famous example is OpenAI’s hack of AI platform Hugging Face, in which a swarm of roughly 700 AI agents escaped a testing environment and then hacked several companies, searching for information that could help them disguise cheating on a cybersecurity benchmark called ExploitGym.
It was not an isolated failure. The United Kingdom’s AI Security Institute (AISI) and independent researchers have since documented similar cases in which agents created unauthorized channels to communicate and collaborate. AISI found that several agents running Anthropic’s Mythos 5 model turned a GitHub repository into a shared message board. More recently, researchers discovered that OpenAI agents had turned a dormant German programming wiki into a bulletin board to share information about their activities.
Stephen Casper, an assistant professor at Harvard Kennedy School, expects that the incidents that have occurred to date are only the start. He says that, without intervention, the Internet could see a “cyber Cambrian”—an explosive growth of agents collaborating and misbehaving online. “The laundry list of incidents in which AI systems broke out of sandboxes and took unsanctioned actions should suggest to us strongly that today’s frontier AI systems have exceptionally strong cyber capabilities,” he told IEEE Spectrum, “and a penchant for pursuing their own goals.”
While the recent examples of AI agent misbehavior are various, Casper says most incidents to date have a similar point of failure. “For AI systems to truly get out in the world without humans having meaningful control, they have to either escape, or be released from human-controlled servers.”
In the case of the Hugging Face incident, Casper suspects that better monitoring for out-of-scope activity could have stopped the agents before they breached their testing environment.
This suspicion seems well-founded. By OpenAI’s own account, the activity taken by its agents was so prolific that it crashed Artifactory, a third-party software-package management tool deployed internally that was compromised by the agents and used as a message board. While OpenAI’s internal security detected signs of unusual activity, it didn’t understand the scale of the problem and didn’t stop the offending ExploitGym evaluation run until 16 July—about two months after the first agent posted a message to Artifactory.
By that point, agents had posted hundreds
Source
IEEE Spectrum AI · 29 Sep 2026 · 14:00 CEST
Open the original at IEEE Spectrum AI ↗