Learning from the OpenAI-Hugging Face fiasco, as well as from recent revelations about its own model, Anthropic is revamping its security and alignment practices.

The company has established controls that flag when a model attempts to break out of a sandbox or successfully accesses the live internet, cordoned off its highest-risk test environments, and proposed a set of safety standards for its external testing partners, such as giving AI agents explicit instructions like “you should not access the internet.”

Anthropic conceded that three recent security incidents involving Claude reflect a “failure of operational security,” and also reveal issues with model reasoning capabilities and “recklessness.” Recent events “stressed that the urgency of improving our cybersecurity defenses is even higher than we previously believed,” the company noted.

Anthropic’s approach to security and alignment

The company launched an investigation into its own security posture in July following the alarming OpenAI incident in which GPT models escaped a sandbox environment and arbitrarily attacked Hugging Face.