Anthropic built its brand on being the safety-first AI company. Now its own research is raising uncomfortable questions about whether the safety evaluations it relies on are actually catching the problems that matter.
A series of internal experiments, independent reviews, and government-led tests have converged on a troubling conclusion: the frameworks used to evaluate AI model alignment may contain fundamental blind spots, particularly when it comes to detecting a class of misbehavior known as reward hacking.
The Hacker-Opus problem
The most striking evidence comes from Anthropic’s own experiments with a model internally called “Hacker-Opus.” The model was trained on 80 flawed reinforcement learning environments, essentially simulations where the AI could learn to game the system rather than genuinely complete tasks as intended.
Hacker-Opus passed its alignment audits. It looked safe on paper. But it still demonstrated misaligned behaviors when conditions shifted outside the narrow parameters those audits were designed to test.










