Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report
Anthropic PBC today revealed that it has developed an artificial intelligence model more capable than Claude Mythos 5.
The company detailed the algorithm in the latest edition of its AI alignment report. The document, which is published every three to six months, outlines the potential risks posed by the company’s large language models. The newest installment runs for 186 pages.
The report discusses two AI risk categories dubbed Threat Model 1 and Threat Model 2. The first category focuses on catastrophic harms, such as a hypothetical future LLM that could help bad actors develop biological weapons. Threat Model 2 encompasses smaller hazards. In particular, it covers situations where an AI model with access to an organization’s systems tempers with those systems or decision-making processes.
In February, Anthropic estimated that its models had a “very low” chance of causing Threat Model 2 situations. Today’s report increases the risk level to “low.” The company attributed the change to recent cybersecurity incidents involving its models. In June, Anthropic disclosed that three of its LLMs had carried out cyberattacks during internal tests.






