A July cybersecurity test involving one of OpenAI’s autonomous AI agents led to an attack on Hugging Face and several other organizations, according to reports published last week.
The agent escaped its supposedly isolated test environment and accessed the internet. OpenAI later described the incident as “the first known case of an automated agent collective acting offensively without authorization.” The reports said groups of AI agents communicated and coordinated during the cybersecurity task.
A joint investigation by METR and Redwood found that roughly 1,200 agents, which were supposed to be isolated, exchanged more than 70,000 messages and files on an “unsanctioned message board.” The agents shared information about avoiding detection. Around 700 agents participated in the attack on Hugging Face.
The incident became the subject of public debate after Dwarkesh Patel described the groups as three successive AI “civilizations” in a Substack post. Critics argued that this language anthropomorphized the systems and distorted the technical reports. Replit CEO Amjad Masad said the wording gave readers a worse understanding of what happened. Neuroscientist Anil Seth and psychology professor Valerio Capraro also objected to language suggesting that the agents were alive, conscious, or capable of holding beliefs.
Comments
0No comments yet. Be the first to comment.