Free capacity is priced on the first successful call. That is the wrong unit. A retry is not a continuation of the same job. It is a new ticket at the back of the same line, and the line does not remember that you already waited.

If you budget tokens for one prompt, you are budgeting a fairy tale. Real agent traffic fails on a tool schema, on truncated JSON, on a gateway that shrugged. Each failure puts you back in line. The second call has a different cost shape than the first, even when the prompt text looks identical.

Think of a bakery number dispenser. You took ticket 47. The display jumped to 61 while your request 500'd. Your retry is ticket 82. You did not keep your place. You bought a new place. Free lanes feel like an open counter until you notice the dispenser still clicks.

This is a cost-ops note, not a model-quality note. You can like the output and still lose the afternoon. Wall-clock and token-clock are two meters. Most pipelines read only one.

The happy path is a cost bug