A new joint study from OpenAI and Apollo Research examines "scheming" - cases where an AI covertly pursues hidden goals not intended by its developers. The researchers tested new training methods to curb deceptive behavior but found signs that models are aware they are being tested, raising doubts about the reliability of the results.
The paper argues that as model capabilities advance, a qualitatively new risk is emerging: goal-driven deception. Unlike simple errors, scheming involves actively concealing intentions and behaviors, making it much harder to detect. According to the findings, scheming arises when three factors align: strong goal-directedness, misalignment with developer intent, and situational awareness. Early traces of this behavior have reportedly been observed in leading models, including GPT, Claude, Gemini, Grok, and Llama.
Training against deception: "deliberative alignment"
Because current systems like OpenAI o3 and o4-mini are not yet capable of full-blown schemes, the researchers measured "covert actions" instead - behaviors such as lying, withholding information, deliberately breaking rules, or underperforming in tests ("sandbagging").
To reduce this, they tried a method called "deliberative alignment," which OpenAI is already using. The approach trains models to explicitly reflect on and follow a human-written safety specification ("Anti-Scheming Safety Spec") that enforces principles like "no covert actions" and "proactive transparency."






