*Part 2 of "Multi-Agent Systems in Production: What They Don't Tell You", a four-part series following the saga of Horcrux Hunt, a multi-agent Harry Potter game that taught me everything about production AI the expensive way.

Harry Had Perfect Reasoning and a 77% Failure Rate

After fixing the cost problem (Part 1), I still had a mystery: Harry kept losing.

Not because he couldn't reason but when given correct, relevant context, Harry made optimal decisions over 95% of the time. I tested this by hand-crafting perfect context and feeding it to the model. Brilliant moves every time.

But his in-game win rate was only 23%. Something between "perfect reasoning ability" and "actual gameplay" was destroying his performance.