A Saturday afternoon alert: under-replicated partitions on the production Kafka cluster, well over the threshold. Status CRITICAL, replicas out of sync, risk of unavailability if a broker fails. The kind of page that makes you put down whatever you were holding.
The cluster is self-managed Kafka running on EKS, which means there is no AWS support engineer on the other side of this — whatever is happening, it's yours to diagnose.
By the time anyone looked at it, the number was zero. Nothing was broken. Nothing needed fixing. The cluster had rotated its own TLS certificates, restarted every broker one at a time exactly as designed, and gone back to green — and the monitor had caught the middle of that perfectly healthy operation and called it an emergency.
This is what that incident actually looked like from the inside, and why a 200-line Kubernetes operator nobody talks about was the reason the blast radius stopped where it did.
The architecture






