CrashLoopBackOff is the one everyone watches. Here are 20 Kubernetes failures — in logs and events — that actually deserve an alert.

It's 3am and the dashboard is green. Checkout is timing out anyway, because CoreDNS quietly fell over twenty minutes ago and nothing you alert on noticed. Most teams running Kubernetes alert on CrashLoopBackOff and OOMKilled and call it covered — that's maybe a third of the failure modes that take down real clusters. The other two-thirds live in pod events and kubelet logs, and they stay invisible until the first incident teaches you they exist.

Here are 20 of them. Each has paged someone. Each shows up in your cluster's logs or events before it shows up on a status page — if you're looking.

Pod lifecycle failures

1. CrashLoopBackOff