Also published on the CNCF blog. Cross-post with canonical link to the CNCF version.
The 3 AM Call
We got paged one Tuesday morning. A critical production service had crashed under traffic—not gradually degraded, but crashed. Hundreds of pending pods. Users were seeing 15–20% error rates. The incident postmortem was brutal: reactive autoscaling had fired, but it was already too late.
The timeline looked like this:
06:00 Traffic spike arrives






