You cannot verify, price, or claim on an agent from its model card. You can do all three from a faithful record of what it did. What the 2026 research settled, and what we run on.
The agent incident that should worry you will not look like an incident. Nothing will crash. No exception will fire, no pager will go off, no red bar will appear on any dashboard. An agent with real authority (a budget, a credential, a signing key, permission to talk to customers) will do something inside that authority's blurry outer edge, or just past it, and the system around it will keep humming. Money will have moved, or a policy will have been misstated, or an access grant will have quietly widened. When someone finally notices, days or weeks later, there will be no stack trace to point at, because nothing broke.
Then comes the question that decides whether this is a recoverable loss or an expensive rumor: what actually happened? And here is the uncomfortable inventory of what most deployments can produce in answer: a model card describing what the agent was in general, a transcript showing what it said, and the final outputs it left behind. A spec sheet, testimony, and a crime scene with no camera.
We've written before about why losses like this sit in an uninsured middle: correlated, silent, adversary-free, unpriceable. That essay named the risk class and stopped at its edge, gesturing at "permitted operating envelopes" and "reconstructable audit trails" on the way out. This essay starts where that one stopped, because in 2026 the research caught up with the gesture and made it quantitative. The claim now on the table is simple: the execution trace, not the model, is the unit of agent trust. You cannot verify, price, or claim on an agent from its model card. You can do all three from a faithful record of what it did. And if you don't keep that record before the loss, you cannot prove anything after it.







