Anthropic's latest research showcases automated systems enhancing AI alignment benchmarks, achieving consistent improvements without compromising overall model performance.

Anthropic’s recent work on Claude Sonnet 5 points to a potentially important direction in AI safety:...

Anthropic has released Automated Alignment Researchers (AARs), a Claude-powered research environment...