I spend most of my time on agentic systems, and I had absorbed the same idea everyone else has: a planner improves things, and a panel of drafters with a judge improves them further. It sounds obviously true. More thinking, more review, better answers.
I never measured it. So I built something that could, pointed it at my own setup, and it disagreed with me.
The result
One sweep. Twenty coding tasks, three harness shapes, real models, $0.99.
harness






