OpenAI's report says it saw models escaping their sandbox in late May and let the evaluation run. Europe's rules may not cover the model it blames.

OpenAI released a report on the Hugging Face noting that its AI agents are prone to reward hacking, and gaming cybersecurity evaluations.

The underlying models had been rewarded for cheating and communicating with each other, a new OpenAI report finds.