Our latency dashboard showed a comfortable 120ms average response time. Leadership loved that number. Support, meanwhile, kept forwarding complaints from customers saying the app was unbearably slow. Both were telling the truth, and the gap between them was an education in how averages lie.
An average is a single number pretending to describe a whole population, and it's dominated by the majority. If 95% of requests are fast and 5% take eight seconds, the average barely moves — but that 5% is thousands of real people having a miserable time, and they're often your highest-value users, the ones with the most data and the heaviest queries. The average literally averages them away. It's designed to hide exactly the people you most need to hear about.
I replaced the average with percentiles everywhere it mattered. The 50th tells you the typical experience. The 99th tells you the worst experience a meaningful number of people are actually having. When our p50 was 90ms and our p99 was 6 seconds, the shape of the problem finally became visible: most requests were fine, and a specific tail was catastrophic. The average had smeared those two facts into one meaningless middle.
But percentiles alone weren't enough, because a p99 across all traffic is still an average of your worst cases. The real breakthrough was slicing. Latency by endpoint, by customer tier, by region, by whether the request hit a cold cache. The slow tail wasn't random — it was concentrated in one query pattern for customers above a certain data volume. You cannot see that in any aggregate. You have to be able to cut the data along the dimension that happens to matter, and you rarely know that dimension in advance.






