I built a hostile AI agent receipt that passed the cryptographic checks.
The signature was valid. Sequence coverage was complete. The evidence still failed to prove the claim.
The correct answer was undecidable.
That failure exposes a gap in how AI agent receipts are being compared. An authentic record is valuable, but authenticity alone does not tell another system whether the evidence is strong enough to authorize what happens next.
I think receipt systems have three separate jobs.






