I had a three-agent setup: a planner, a researcher, and a critic passing work between them. I graded each one carefully. The planner made clean plans. The researcher cited its sources correctly. The critic caught weak claims. Every agent scored around 0.9 on its own tests, and I was happy.

The team was still wrong about a third of the time.

The plans were good. The research was good. The critiques were good. But somewhere between one agent and the next, things fell apart. The planner would say "research scaling laws, but skip this one paper, the user already has it," and the researcher would cite that exact paper two turns later. The constraint just vanished in the handoff. Nobody's individual turn failed a test, and the final answer was wrong anyway.

That is when it clicked: I was grading the agents when I should have been grading what happens between them.

Multi-agent is not single-agent times three