Anthropic just published research showing its Claude models can autonomously identify and fix AI alignment failures better than human safety researchers can. The paper, titled “Automated researchers can reliably mitigate alignment failures,” describes systems that improved performance across ten distinct categories of misaligned AI behavior without degrading the models’ general capabilities.
What the automated alignment researchers actually do
Anthropic’s Claude models now function as what the company calls automated alignment researchers, or AARs. These systems autonomously devise, assess, and enhance methodologies designed to mitigate specific categories of AI misalignment, including privacy violations and deception.
The results were tested against public benchmarks covering ten failure categories. Every single one showed improvement. The methods that worked best also generalized to held-out benchmarks and the open-source Petri auditing tool, scenarios the system wasn’t specifically optimized for. The techniques even remained effective on models up to 4.7 times larger than the ones originally used for optimization.
Claude achieved roughly 85% gap closure on deception benchmarks. When 28 human safety researchers were given the same alignment tasks under comparable constrained conditions, Claude’s automated approach delivered 20% better performance on deception tasks than the strongest human proposals.








