Part 3 of **The Answerability Problem. Part 1 showed the standard harness excluding the questions that test refusal, and my own system scoring 0.000 on them. Part 2 measured six candidate signals and found none that beat plain cosine. This part explains why both were asking a question that has no answer. Code and every number: RE-call (MIT).
Two public benchmarks disagree about whether memory systems can abstain.
On LOCOMO's adversarial category my system scores 0.000 out of 446, and every lever I ship reaches a usable catch rate only by refusing 60% of legitimate questions. On BEAM's abstention category the same class of system does something, and the incumbent's own published answers abstain correctly 38 times out of 70.
For weeks I treated that as two data points about one capability and tried to reconcile them. It is not two data points. It is two samples of a hidden variable, each reported as a scalar.
Name the variable and the disagreement dissolves.






