Most "AI pentester" projects are a single LLM in a while-loop with a shell. You

give it a target, it runs commands until it decides it found something. That's

how you get confident nonsense — a model that writes a beautiful vulnerability

report for a bug that doesn't exist.

I wanted the opposite: an engine where a finding has to be earned. So I built