Part 1 ended with a confession: you'll never test your way to 100% correctness. For critical workflows the final safety net is a human. But "human in the loop" has a dirty secret. Badly designed, it's theater.

The rubber-stamp problem

Route every agent action to a human for approval and watch what happens. Week one, careful reviews. By week three approval fatigue sets in and people click approve at the speed of thought. You've paid a human salary to become an Enter key, and the loop provides zero actual oversight.

The goal isn't humans reviewing everything. It's humans reviewing exactly the things that need judgment, at a volume they can sustain.

Designing oversight people can actually do