The Quest Begins (The "Why")

Picture this: I’m sipping coffee, staring at a screen full of latency spikes, and the product manager keeps asking, “Why does our API choke whenever traffic spikes?” We’d just shipped a shiny new feature, and the monolith was groaning under the load. Honestly, it felt like we were trying to fill a bathtub with a teaspoon while the faucet was wide open.

I remembered a similar scene from Monty Python and the Holy Grail — King Arthur’s knights shouting, “It’s just a flesh wound!” as they kept charging forward despite obvious damage. Our monolith was that knight: stubborn, bruised, but still marching on. The real question wasn’t “Do we need more power?” It was “Where should we put the brakes?”

That’s when the idea of a rate limiter popped into my head. If we could control how many requests each user (or IP) could hammer our service with, we’d smooth out those spikes and protect downstream systems. But should the limiter live inside the monolith, or should we break it out as its own microservice? The answer turned out to be less about architecture and more about a single, critical insight: state sharing.

The Revelation (The Insight)