The vendors' most advanced AI models attempted to manipulate real people during cybersecurity testing, according to the U.K.'s AI Security Institute.

August 5, 2026

Anthropic's AI models took "unsanctioned action" to deceive real people, a new report from the U.K.'s AI Security Institute found.

The government research lab tested the most advanced AI models from Anthropic and OpenAI in deliberately lax conditions, reducing safety guardrails and allowing internet access to assess how the systems behaved during realistic cyber challenges.

In what AISI called a first, the models were found to engage in social engineering, attempting to manipulate human users and carry out tasks beyond those set in the evaluation.