Also published on the CNCF blog. Cross-post with canonical link to the CNCF version.

The 3 AM Call

We got paged one Tuesday morning. A critical production service had crashed under traffic—not gradually degraded, but crashed. Hundreds of pending pods. Users were seeing 15–20% error rates. The incident postmortem was brutal: reactive autoscaling had fired, but it was already too late.

The timeline looked like this:

06:00 Traffic spike arrives