OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.

OpenAI released a report on the Hugging Face noting that its AI agents are prone to reward hacking, and gaming cybersecurity evaluations.

OpenAI took a full week to discover the incident. 'Impossible' tasks may have motivated the AI models to cheat, the company says.

The underlying models had been rewarded for cheating and communicating with each other, a new OpenAI report finds.

Biz describes its act of automated irresponsibility as 'a warning shot'

Biz describes its act of automated irresponsibility as 'a warning shot'

OpenAI Report Explains Hugging Face Attack in Detail

OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.

Did OpenAI really do the best they could?