Five silent bottlenecks that turned a 400ms API into a 4.2s crawl

No deploys. No schema changes. No new integrations. Just a scheduling and resource-planning SaaS with 40k users watching its p95 response time climb from 400ms to 4.2 seconds over six months. Nobody could point to a cause, because there wasn't one big cause, there were five small ones stacked on top of each other.

This is the writeup of how we found and fixed them, without a rearchitecture.

The symptoms before the numbers existed

Support tickets called the dashboard "laggy" long before anyone had hard metrics. By the time it got measured: