I look after a platform of 40-plus production databases in healthcare. Multi-terabyte, HIPAA-scoped, 99.9%+ availability targets, and a rotation of engineers who carry the pager for it. When I joined, the team was competent and the systems were reasonably instrumented. We still had a repeat-incident problem.

Not a volume problem. A repeat problem. The same class of page kept coming back, sometimes to a different engineer, often a week or two apart, and each time it got handled fresh. Somebody would work it at 3 AM, apply the mitigation that made the symptom stop, close the incident, and go back to sleep. Entirely reasonable behavior. Also the exact mechanism by which an organization pays for the same outage four times.

The instinct in that situation is to reach for tooling. Better alert correlation, a new observability tier, an incident platform with nicer timelines. We had Datadog and Grafana and they were fine. Tooling was not the constraint.

The constraint was that context died at the handoff.

What a readiness review actually is