When you ask a language model why a service broke it will tell you. Confidently, fluently, and with assurance, every time, regardless of whether the evidence supports the conclusion it reached or not.

For a post-incident review, that is worse than useless. A plausible answer with nothing behind it still gets written into the postmortem, and somebody WILL act on it.

Because of this issue, I built the chain the other way around. It is computed first, deterministically, and a model is only allowed to narrate what has already been established. Below is the design, the specific numbers it runs on, and what it prints.

Three sources that never mention each other

It is the morning after an incident. Your checkout API began failing readiness probes at 03:20AM and as the DevOps, you need a timeline in order to create the review.