Anthropic’s recent work on Claude Sonnet 5 points to a potentially important direction in AI safety:...

Anthropic’s recent work on Claude Sonnet 5 points to a potentially important direction in AI safety:...

Anthropic's Claude-powered automated researchers closed 26% to 96% of the safety gap across 10 alignment failures, outperforming human