I build a hospital's internal tools as a non-developer, which means AI writes most of the code. So of course I did the responsible thing: I had AI review it too.

Not casually. I set up separate roles — one agent implements, another runs the tests, another reads the code specifically for security problems. Different instructions, different focus, different pass over the same work. It felt like a review process. For months, the security reviewer flagged almost nothing.

I read that as good news.

Then, over three days on this site, strangers found eight real defects in the same systems. Not one of them came from my reviewer.

The number that made me stop