Part 2 of **The Answerability Problem. Part 1 showed the standard harness excluding the questions that test refusal, my own system scoring 0.000, and every lever I ship failing. It asked whether any cheap signal can separate "relevant" from "answerable." This is the bake-off. Code: RE-call.

LongMemEval is the other public benchmark in this space, and unlike LOCOMO it names abstention as a first-class question type. 500 questions over a synthetic user's chat history: 470 answerable, 30 whose answer was never mentioned.

My retrieval does well on it. The right session is in the returned set 97% of the time.

Then the trust layer refuses to answer 48% of those.

Sit with the shape of that. The system finds the answer, holds it, and declines to use it, in nearly half of the cases where it succeeded. As a product that is worse than having no abstention layer at all: you have paid for retrieval quality and then thrown it away at the last gate.