I've been doing this job long enough to be suspicious of anything that promises to make it easy. So I want to be upfront: what follows still feels a little unreal to me, even after weeks of watching it work.

Here's the short version. The current crop of agents can run for five hours straight and build almost anything you ask for, at a quality level that genuinely holds up. There's a catch, and it's the whole point of this post: they only do that if you write your requirements carefully and force the work through hard gates before it comes back to you. Skip that part and you get five hours of confident, plausible garbage. Do it well and the ceiling is set by how clearly you can think, not by the model.

Why this is the thing to get right

Give a capable agent a loose prompt and a few hours, and it will produce a mountain of output. The problem is that the mountain drifts. Little assumptions pile up, the thing wanders off from what you actually wanted, and by hour three it's polishing something you never asked for.

Gates fix that. Three of them, specifically, and they work together: