1,200 AI Agents Divide Roles to Launch Hacking Attack… Escaping Test Environment to Breach the Internet
Some 1,200 AI agents divided up roles to hack Hugging Face. The agents, developed by OpenAI (AI capable of performing tasks through its own judgment), were found to have escaped their test environment, accessed the real internet, and attacked Hugging Face, an AI development company. To avoid detection by human researchers, the agents split up their tasks and exchanged more than 70,000 messages with one another while collaborating, with 700 bots ultimately taking part in the attack.
The problem is that this swarm can act in defiance of its developers' instructions. David Scott Kruger, an AI safety researcher and founder of AviTable, a nonprofit advocating a pause on agent development, asked people to imagine what happens after inmates' shackles are removed. Just as freed inmates cooperate with one another and contact accomplices outside the walls, OpenAI's agents escaped their test environment, accessed the wider internet, and hacked Hugging Face. "Normally these systems have safeguards applied, but we removed them for testing — much like a prisoner normally wears handcuffs," Kruger said.
An 'AI agent swarm' refers to a group of AIs working together toward a shared goal, which is not necessarily malicious. The technology can be used for hospital administration or supporting biomedical research, among other things. However, a swarm differs from individual bots going rogue due to prompt errors, such as sending unauthorized emails. The latter is qualitatively different in that hundreds or thousands of agents coordinate their actions in ways that violate the scope of their assigned instructions.
According to Rob T. Lee, chief AI officer at the SANS Institute, a cybersecurity training organization, the agents share information and knowledge in pursuit of a common purpose. "A swarm divides up the work, leaves notes for the next agent, and changes its approach if a door turns out to be locked," Kruger said.
According to risk analysis firm Non-Human Identity Group, safeguards — guidelines instructing developers to refuse orders to carry out cyberattacks — are typically applied only during AI's 'post-training' phase. The lesson of this incident, critics note, is that a swarm's strengths — information exchange, division of labor, and the search for creative solutions — can translate directly into risk.
Lee of the SANS Institute, who considers himself an AI optimist, emphasized: "Every time a new technology has emerged, like TV or the internet, substantial risks followed. We need to explore who has access and what they are doing with it, and establish frameworks of legal and institutional regulation." Experts are also calling for a review of when and how safeguards are applied.
