Almost every productivity figure attached to an AI rollout comes from asking people how much time they saved. That method has a predictable direction of error, and the fix is not a better survey — it is a design in which the answer is not supplied by the people being measured.

Narrowing “productivity” to something countable

Productivity is output per unit of input. Before anything can be measured, both halves have to be fixed to something a query can count, and the narrowing is where most of the intellectual work is.

Pick a task, not a role. “Engineering productivity” is not measurable in a quarter; “median time from pull request opened to first review comment” is. “Legal team efficiency” is not measurable; “contracts reviewed per reviewer-day, for contracts of type X” is. The narrowing feels like a retreat and it is the opposite: a narrow metric with a real baseline supports a conclusion, and a broad one supports a slide.

Then fix the input side, because this is where the arithmetic quietly breaks. Time saved is only a gain if the time went somewhere countable. If drafting takes twenty minutes less and the person spends those twenty minutes on the next ticket, output rises and you can see it. If they spend it in a meeting that was going to expand anyway, nothing observable changed. Both are real outcomes; only one is a productivity gain, and a design that cannot tell them apart will report the first while measuring the second.