The metric that quietly decides your agent bill isn't the input price of your model. It's how much you pay to read what you already sent.
If you run multi-step agents, you've probably had this moment: a task you expected to cost pennies comes back as a small surprise on the invoice. You didn't change models. You didn't prompt more. So where did the tokens go?
Most of the time, they went to paying for the same context, again and again.
A 2-minute task can fire 40+ billable calls
Here's a shape I keep seeing. An agent does a "2-minute" job:






