Most "AI pentester" projects are a single LLM in a while-loop with a shell. You
give it a target, it runs commands until it decides it found something. That's
how you get confident nonsense — a model that writes a beautiful vulnerability
report for a bug that doesn't exist.
I wanted the opposite: an engine where a finding has to be earned. So I built






