Anthropic's Claude-powered automated researchers closed 26% to 96% of the safety gap across 10 alignment failures, outperforming human

We had Claude autonomously train models to improve their performance on several public benchmarks that measure 10 categories of alignment failure. For all 10, Claude found fixes…

Anthropic’s recent work on Claude Sonnet 5 points to a potentially important direction in AI safety:...