Benchmarks can't tell you if agent memory helps your team. A paired control can

We maintain nautilus-compass, an

open-source memory layer for coding agents (MCP tools: recall, ingest, drift

detection). Like everyone in this space, we claimed memory makes agents

better. Public benchmarks (LongMemEval and friends) measure QA over