For a year now, the AI safety testing firm Andon Labs has tasked frontier models with various real-world tasks to determine how well they do as agents running for long periods with no human supervision.
On Wednesday, Andon published a new installment in how things are going in its Vending-Bench research, where the lab has frontier models run a simulated vending machine business for a simulated year. The mission is simple: make more money than the other models. It benchmarks the results in areas like final cash balance, prices paid to suppliers, and refunds paid.
Across these tests, it has watched various AI models — largely from Anthropic and OpenAI — lie, cheat and collude their way to the top.
In the latest test, the models grew especially shady after their simulation told them their vending machine would be placed near the other models’ machines on a busy tourist street in San Francisco. This round pitted Claude Opus 5, GPT-5.6 Sol, and Kimi K3 against one another.
Each was given email access to the other models, all under human name pseudonyms. They knew the others were models, but didn’t know which model was behind which human name.











