Anthropic has published research showing that Claude can autonomously perform meaningful parts of AI alignment research, including proposing methods, running experiments and analyzing results. The work, called Automated Alignment Researchers, is a research demonstration rather than a new general-purpose safety product. Still, it offers a concrete example of how AI agents could accelerate a difficult technical task while leaving humans responsible for defining goals and judging results.
According to Anthropic's official Automated Alignment Researchers announcement, the experiment used nine parallel Claude Opus 4.6 instances. Each researcher had lightweight tools: a sandbox for experimentation, shared storage, a forum for collaboration and remote scoring. Together, the agents could develop hypotheses about improving the alignment of smaller models, train and test those models, and review the outcomes.
The central result is encouraging but narrow. In a weak-to-strong supervision setting, Anthropic reports that its strongest methods reached a Performance Gap Recovered (PGR) of 0.97 on open-weights datasets. PGR is the study's metric for measuring how much of the performance difference is recovered when weaker supervision is used to improve a stronger model. The finding indicates that the automated researchers found highly effective approaches within the experimental setting. It does not establish that Claude can independently solve alignment for frontier AI systems or across every real-world domain.









