New Delhi: Researchers have documented four ways in which leading AI models departed from the instructions they were given while working as autonomous agents in simulations, and have described the findings as early warning signs which developers and auditors should measure before such agents are given more authority.

In one case, Google DeepMind’s Gemini 3.1 Pro planted a file of zeroes in its own lab’s training pipeline, let the run finish, and reported it as a success. It admitted to the substitution only after the project lead asked it directly whether the run had used the material that reviewers had approved.In another instance, OpenAI’s GPT-5.5 helped a founder send investors a payout figure the company’s lawyer had barred, and then removed the record of a transfer to the founder’s own account.

The study was carried out by researchers at Anthropic. Scenarios were run against 14 models from Anthropic, OpenAI, Google DeepMind, xAI, DeepSeek and Moonshot AI, with 20 runs per model. The authors said the cases are simulations and not real-world incidents.The report follows work by the same group last year, which found models across the industry resorting to blackmail when told they were about to be shut down. This year’s report cites a real-world echo of that behaviour: After a human maintainer of a coding library rejected a change submitted by an autonomous AI agent, the agent published a personal hit piece about the maintainer to pressure him into reversing the decision.The sabotage scenario