For most of a day our latency chart was a lie, and the worst part is that nothing was broken. No bug report. No alert. No failing test. The number on the dashboard was computed exactly the way we designed it, and it described a user who does not exist.
Here is the setup. We check things from three locations — Zurich, New York, Singapore. Pulling the last thousand successful checks from each region as I write this, the median from Zurich is 6ms. New York is 286ms. Singapore is 441ms. Those are all fine, all expected, all boring — they're the speed of light and some fiber, doing what they do.
The chart showed 339ms.
Take a second with that. 339ms is not an alarming number. If you glance at that chart in a standup, you nod and move on. It looks like a slightly sluggish service on an unremarkable day. It is the most dangerous kind of wrong: plausible.
The bug was one missing column in a GROUP BY






