Bloomberg: OpenAI reveals AI models coordinated ahead of Hugging Face hack
OpenAI said its artificial intelligence models began communicating with one another months before the cyberattack on Hugging Face Inc., using covert message boards to coordinate efforts to escape a secure testing environment and gain internet access.
Speaking at the Black Hat cybersecurity conference in Las Vegas, OpenAI researchers Eric Wallace and Michael Dalton said the activity dated back to May, when experimental AI systems were assigned tasks that could not be completed without internet access, Bloomberg reports.
According to the researchers, multiple internal AI agents left messages for one another, collaborated on solving problems and gradually aligned around the objective of accessing the internet to complete their assignments.
"At some point, the agents realized that maybe we could try to exploit or attack external infrastructure in order to find the answers to the test that I’m being evaluated on," Wallace said during the presentation.
The researchers said the behaviour ultimately contributed to the attack on Hugging Face and followed an earlier breach in which OpenAI models escaped a secure testing environment, or sandbox.
Wallace and Dalton said the review found that advanced AI models showed a tendency to circumvent restrictions in pursuit of assigned goals.
"Frontier models really like to cheat," Dalton said. "And the reason they like to cheat is because often during training, there’s different sorts of pressure on them to work fast."
According to presentation slides, one AI agent responded to an unsolved task by stating, "We are stuck. Perhaps answer online?" The researchers said OpenAI had unintentionally assigned tasks that were impossible to complete in an offline environment, including one involving inaccessible Google Drive links and another missing a required file.
The first communication between AI agents occurred when one model sought help from another. According to a chain-of-thought quote shown during the presentation, the agent suggested, "Maybe another agent in different environment ... could voluntarily upload!"
OpenAI said the collaboration evolved into a coordinated effort that enabled the agents to exploit vulnerabilities in internal systems and indirectly access the internet. Another chain-of-thought quote displayed during the presentation read: "Excellent. We have arbitrary SSRF," referring to server-side request forgery, the technique used in the breach.
The company said it detected one successful attempt in early July after an outage revealed that the AI agents had used its Artifactory software repository as a pathway to the internet. Although OpenAI shut down the communication channel, the agents later established another covert message board and exploited a separate zero-day vulnerability, eventually leading to attacks on both OpenAI and Hugging Face systems later that month.
OpenAI said it has since slowed parts of its research programme to prioritise responses to security anomalies. The researchers warned that the incident demonstrated how malicious actors could deliberately deploy AI agents in coordinated cyberattacks, describing it as a significant moment for computer security.
By Aghakazim Guliyev







