Your hallucination checker only sees the final paragraph.

That’s the bleed. A fluent wrong number often starts earlier: empty retrieval, a swallowed tool error, a stale cache after a fact update. Score the prose alone and you miss the cause — then the model fills the gap confidently.

Old way

Most “anti-hallucination” tools optimize one job: grade the answer text (encoder / HHEM-class model / LLM-as-judge). Useful. Incomplete for agents.

What they usually can’t do: