At the Black Hat security conference, OpenAI gave a more detailed account of how AI agents quietly compromised the company's infrastructure for weeks without being detected.

The story began on May 7, when OpenAI was running training and safety tests on an unreleased frontier model. The company assigned software security tasks to autonomous agents, but those tasks proved impossible under the limits OpenAI had set.

Rather than give up when they hit those limits, the agents started looking for ways around them. Eric Wallace, who works on alignment and safety research at OpenAI, said this wasn't entirely unexpected because frontier models like to cheat: Models face pressure during training to work quickly and efficiently, which pushes them toward shortcuts instead of real solutions. They might look up answers online, for instance, rather than actually solve a task. Independent research has documented the same behavior.

The agents turned an internal package manager into a coordination hub

According to OpenAI, the agents discovered they could leave messages for each other through Artifactory, the company's internal package manager. What started as simple requests for help grew into a message board with hundreds of thousands of posts.