Your agent passed the eval, so you shipped. The next day a user sends almost the same input and it fails. Nothing changed. You just learned that "it passed" was one sample of a distribution, and you shipped on a coin flip that landed heads.
Part 1 defined the bar. Part 2 built the eval set that measures it. Part 3 is the uncomfortable part: an agent will not give the same answer twice, so a single green run tells you almost nothing about AI agent reliability.
One run measures luck, not reliability
A deterministic function passes a test or it does not, forever. An agent is a random variable. Run the same case ten times and you might see nine passes and one confident wrong answer. A single run showing the pass is not evidence the agent is good, it is evidence you got lucky once. Reliability is not a checkmark. It is a distribution, and you have to measure its shape.
Measure the distribution, not one draw






