Co-authored with Netanel Kadosh

You've built a great inference service. Auto-scaling is configured. The model performs beautifully in testing. You ship it to production feeling good.

Then a traffic spike hits at 2 AM.

New pods start spinning up. Users start waiting. Your on-call phone lights up. You watch the dashboards- nodes are healthy, no errors, but the pods just... aren't ready. Eight minutes pass. Then nine. Then ten. Sometimes fourteen.

Fourteen minutes for a service that responds in under 200 milliseconds once it's warm.