A manual checkpoint outperformed full automation by 35 percentage points. That's the number that changed how I build every prompt chain now.
I spent two months convinced longer chains meant better output. More refinement steps, closer to correct. So I built a 7-step chain for client ad copy — brief intake, angle extraction, brand voice filter, headline drafting, scoring, rewriting, final polish — all automated via Claude Sonnet, each output feeding the next. It worked for two days. Then a client brief came in with an ambiguous audience definition. The angle-extraction step produced garbage, everything downstream inherited it, and because nothing interrupted the chain, I didn't catch the failure until step 6. Burned tokens, burned time, nothing usable.
The fix wasn't a better prompt. It was a shorter chain with me reading the output in the middle.
I rebuilt it as two calls with a manual gate between them. Step one extracts structure — audience, offer, constraints, tone. I read that output. If something's off, I edit it in place, which takes about 30 seconds. Then step two gets the corrected structure and generates headline variants and a body draft. I tracked this across 11 client accounts for six weeks. Usable first drafts went from roughly 40% to around 75%. The gate was the entire reason — not the prompts themselves.






