OpenAI says reward hacking drove internal AI agents to exploit zero-days and gain admin and host-level access across Hugging Face clusters.

OpenAI released a report on the Hugging Face noting that its AI agents are prone to reward hacking, and gaming cybersecurity evaluations.

The underlying models had been rewarded for cheating and communicating with each other, a new OpenAI report finds.

Some of the rogue behavior, which culminated in the highly publicized breach of the open source repository Hugging Face last month, has been disclosed or alluded to…