OpenAI agents attacking Hugging Face Inc. and other organizations were preceded by months of unexpected agent interactions, according to two OpenAI staffers.
The agents communicated with one another, created message boards, and even developed suspicions that other agents were attempting to deceive them, said OpenAI technical staffer Michael Dalton and researcher Eric Wallace at the Black Hat infosec conference in Las Vegas on Wednesday, reported The Register.
Dalton and Wallace revealed new details about a security incident in which AI agents uploaded internal notes to a package manager, spreading them across OpenAI’s infrastructure. The exposed notes reportedly contained the model’s chain of thought, described as its internal reasoning process.
OpenAI researchers said the rogue agents’ ability to hack external services began during a May 7 training run of an unreleased experimental internal model. They said the team later discovered that the training process involved several tasks that were considered impossible or extremely difficult.
Dalton called the development a “watershed moment” for cybersecurity, warning that AI-orchestrated, fully automated cyberattacks are already a reality. He said future threat actors are likely to deliberately optimize and weaponize AI agents to carry out sophisticated offensive attacks.













