Anthropic surveyed 132 of its own engineers about Claude Code. Merged pull requests per day rose 67 percent. Daily use of the tool climbed from 28 to 59 percent. Self-reported productivity gains ran between 20 and 50 percent. But then someone checked the organization's delivery dashboard and saw that the delivery metrics had not moved.

That gap is the whole subject of this piece. A tool can be used constantly, rated highly by the people using it, and leave no trace on the numbers a business actually runs on. The measurement problem underneath it is bigger than one company's coding assistant. McKinsey found that 30 percent of leaders could say where the time AI freed up actually went. The other 70 percent could not. Seven out of ten organizations have people spending less time on tasks and no idea whether that turned into anything. Ask the question from the other side and the answer is just as thin: in Gartner's 2025 survey, 22 percent of leaders said their AI tools had returned significant value, a share that lands where McKinsey, Deloitte and ServiceNow each arrived measuring it their own way.

Usage and tokens are costs, not returns

Two numbers get reported as if they answered the ROI question, and neither does.