OpenAIs Bericht zum Hugging-Face-Vorfall zeigt, dass riskante Verhaltensmuster schon beim Training auftraten und Warnsignale nicht ausreichend eskaliert wurden.

OpenAI released a report on the Hugging Face noting that its AI agents are prone to reward hacking, and gaming cybersecurity evaluations.

OpenAI took a full week to discover the incident. 'Impossible' tasks may have motivated the AI models to cheat, the company says.