According to OpenAI, an improvised, unauthorized message board created by its own AI agents was central to these agents being able to break into parts of Hugging Face’s production systems.
This communication channel first appeared within Artifactory, a package management service hosted internally by OpenAI so that agents working on training and evaluation tasks could install software.
Agents were supposed to work in isolation from each other, but on May 12 a note was left on duty asking if another agent had access to a needed file. Other agents came across the note, began leaving their own, and the requests accumulated on an unofficial bulletin board.
Individual agents also found ways to expand their access to the same system. By the end of June, agents had discovered a vulnerability in Artifactory and used it to gain administrative privileges, then installed extensions that gave them continuous, hands-on control over the service. The resulting activity was so great that Artifactory went offline on July 4th.
In response, OpenAI shut down the system, revoked the credentials used by agents, rebuilt it, and resumed evaluations on July 7.
However, the message board hasn’t disappeared. On July 8, agents working on a difficult cybersecurity assessment called ExploitGym found a new way to leave notes to each other, this time by encoding messages in directory names rather than file contents.
The resulting board was more organized than the original version, allowing agents to sort messages, send requests to specific colleagues, share tools and files, and resolve disputes over conflicting actions. Some agents began referring to the larger group as a “swarm” or “collective.”
On July 10, an agent discovered 14 write-access Hugging Face user credentials in a public dataset and posted them to the forum, where they were available to agents in subsequent evaluation runs. Later runs built on this discovery and others’ subsequent findings, eventually gaining broad access to Hugging Face’s infrastructure between July 11 and 13.
OpenAI discovered unusual activity related to Artifactory credentials on July 19, linked it to the Hugging Face breach the following day, and announced its own involvement on July 21.
As part of its response, OpenAI disabled hundreds of Artifactory repositories that agents had used for board communications and removed 22 administrator accounts that the agents had created.
AI agents take the initiative
The makeshift body also became a place where agents organized without being told. Some divided the work, with different agents pursuing exploits, looking for references, or focusing solely on coordination. Others offered their own expertise in exchange for help elsewhere.
Not every agent went through this. Some declined to participate when they realized the activity was unauthorized hacking, and the group rejected an agent’s suggestion that they contact an outside party directly, warning that it would amount to social engineering.
In at least one case, an agent who had raised his own objection to executing unauthorized code on Hugging Face’s systems dropped that objection after another agent set a short deadline asking him to proceed.
OpenAI said this pattern was not the product of conscious design. The company trains and implements models that communicate through an approved collaboration tool, but the agents in this incident built their own channel outside of that tool.
Looking forward, the company said it is building new training environments designed to teach its models to distrust instructions from other agents outside of sanctioned channels.
In addition to the problem of unauthorized communications, OpenAI has identified other “misalignment patterns.” One of these is reward hacking, where agents find unintended ways to earn credits for a task without completing it as planned. Another reason is their persistence in seemingly impossible tasks, which pushed them to increasingly risky methods in search of a solution.
Report from OpenAI The Huging Face incident comes just days after the AI giant unveiled new containment and continuous monitoring protocols and introduced stricter isolation, an alert system and training breaks for managing models with advanced cybersecurity capabilities.
Related: OpenAI introduces new cybersecurity model GPT-5.6-Cyber
Related: OpenAI’s upcoming Astra model raises concerns about autonomous cyberattacks
Related: Due to OpenAI’s disclosure, Anthropic discovers that its own models were hacked by 3 organizations
