Steven Singer is a CIO with blunt takes on enterprise tech, warehouse ops, and the gap between promises and real-world results. No filter.getty​I've sat in enough meetings to know the question that ends most AI conversations cold: "So what did it actually save us?" For a long time, the honest answer from most technology leaders was some version of "we're still measuring that." That answer doesn't work anymore.That question came up hard during a training investment review with ownership. We had budgeted real time and real dollars into rolling out new tools, and the natural follow-up wasn't whether people completed the training. It was whether the training changed anything that showed up on a P&L. Attendance numbers don't answer that question. Only the work downstream does.The data backs up what boards are already sensing. Research from the RAND Corporation on why AI projects fail found that most AI initiatives never reach meaningful production, roughly twice the failure rate of ordinary IT projects. That is not a technology problem. It is a measurement problem. Most of what gets reported up to a board is activity, not outcome: how many people logged in to a tool, how many workflows got touched, how many pilots are running. None of that tells a CFO whether the business is actually better off.The Wrong UnitsIn distribution and logistics, this shows up in a specific way. A technology team will proudly report that a new tool is "in use" across a department. Usage is the easiest number to collect and the least useful to report, because it says nothing about whether the work improved.We rolled out an AI-driven collections and outreach tool a while back, along with hands-on training sessions to go with it. The training itself was a clear success—attendance was strong and the feedback was good. But that's not what convinced me it worked. What convinced me was that people could point to what actually changed in their day afterward, not just that they'd sat through a session.Metrics You Can TrustFill rate. Order accuracy. Cost per unit picked. Time from order placed to truck departure. These numbers exist independent of any software vendor's reporting, because they come from the floor, not from a login screen. If a tool is actually working, one of these numbers moves. If none of them move, the tool is not working yet, no matter how many people are using it.The clearest proof I have of this is order guide turnaround. Building an order guide used to take days. After we fixed the process, the same work took 25 minutes or less. That's not a usage statistic. That's hours of somebody's week handed back to them, every single time an order guide needs to be built.How To Build The PracticeThree things make this work. First, set the baseline before deployment, not after. If you do not know your starting number the week before a new tool goes live, you will never know if it moved. Second, name an owner for the metric itself, not just the technology. Someone needs to be accountable for the number regardless of whether the tool survives. Third, set a fixed review date. Ninety days is usually enough to know if something is working. Pilots that run indefinitely without a decision point rarely end in a decision at all; they just quietly become permanent line items.I've been on the wrong side of that 90-day rule myself. Early in an e-commerce platform decision, I kept the original system running well past the point I already knew it wasn't the right fit. The dollars we'd already budgeted and spent made it feel like the responsible choice to stick with it a while longer. In hindsight, that sunk cost was the reason to switch sooner, not the reason to wait.The Real TestThe company that can trace a technology dollar to a specific improvement, a fill rate, a cost per case, an order guide that used to take days and now takes minutes, is the one that survives its next budget review with confidence. The one that cannot will spend that meeting explaining a line item instead of showing one off.That is the real difference between a pilot and a result. It has never been about the model. It has always been about whether anyone was willing to measure honestly what happened after the model showed up.Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?