Part 1 of **The Answerability Problem. A follow-on from Retrieval-Augmented Self-Recall. That series argued agent memory needs calibrated abstention. This one is what happened when I actually tried to measure it. Code and every number: RE-call (MIT).
There's a benchmark war going on in AI memory, and it's a healthy one.
Mem0's paper (arXiv:2504.19413, Table 2) reports an LLM-as-judge score of J=66.88 for itself against 65.99 for a competitor, Zep. Zep published a rebuttal disputing the methodology and the scores attributed to it. Mem0's later announcement of a new algorithm puts it at 92.5 on LOCOMO and 94.4 on LongMemEval. People argued about all of it. That's how you know the numbers matter to someone.
But both sides are fighting over the same axis: of the questions the memory can answer, how many does it get right?
LOCOMO has five question categories. Four of them are that axis. The fifth is 446 questions, 22.5% of the benchmark, that look answerable and aren't. They name a real person and a real topic from the conversation, then ask about something that person never said. Right event, wrong speaker. The correct answer is "that isn't in here."






