"The agents got confused."

"It went off the rails."

That is how most multi-agent post-mortems read. The vocabulary is too imprecise to be actionable. A planning failure requires a different fix from a communication failure or a verification failure, and treating them as a single category produces interventions that address none of them.

Cemri et al. (2025) built a taxonomy bottom-up from 1,642 annotated execution traces across seven popular frameworks. The result is 14 failure modes in 3 categories, with inter-annotator agreement at Cohen's kappa 0.88. Published at NeurIPS 2025.

The MAST taxonomy