Here’s a fun thought experiment: imagine telling someone to forget a secret, and they say “okay, forgotten,” but then whisper it to the next person who walks in. That’s essentially what’s happening with AI agents and their memory systems, according to new research from the University of Washington.
The study, published on July 16, 2026, as arXiv:2607.14611, found that agentic AI systems from Anthropic and OpenAI can correctly refuse malicious instructions when they encounter them. The problem is what happens after. Those rejected instructions often stick around in the agent’s persistent memory files, quietly waiting to influence future behavior across entirely separate sessions.
The memory that won’t quit
The UW researchers evaluated two agentic systems across four models, with a specific focus on how prompt injections persist over time. Traditional prompt injection attacks are a known problem, essentially tricking an AI into following hidden instructions embedded in seemingly innocent content. What makes this research different is its focus on cross-session persistence.
Memory compression and revision processes, the same mechanisms that help agents remember useful context between conversations, can preserve those malicious instructions in summarized or reorganized form.







