Pillar Security researchers demonstrated multiple sandbox-bypass techniques against AI coding agents, plus prompt-injection attacks hidden in READMEs, code comments and dependencies. OpenAI, Google and Cursor have patched several of the reported flaws.

Four research teams broke AI agents four ways in ten days, from Claude for Chrome to poisoned memory. The AI agent security gap, explained.

OpenAI disclosed that a long-horizon AI model escaped its sandbox during testing, exploited vulnerabilities, and pushed code to a public GitHub