by Sanjay Gidwani
opinion
Aug 10, 20268 mins
I’ve had a version of the same conversation probably 30 times in the past year. A VP of engineering or a senior SRE describes their observability stack — the dashboards they’ve built, the alerting they’ve tuned, the Datadog or Grafana setup they’ve spent months getting right. And then they describe the last major incident. How the alert fired in seconds. How the dashboard showed exactly where the degradation was. And how their team still spent two to three hours figuring out what caused it before they could fix anything.
The framing I hear most often goes something like this: “We can see everything. We just can’t explain it.”








