If you point any LLM at a target and ask it to "write a security report," it will confidently invent findings that aren't there: an imagined TLS weakness, a "likely" SQL injection, a secret leak that never happened. For a security tool, a hallucinated finding is the worst possible output — you can't hand a client a report built on fiction.

There's a second problem, and it's a hard blocker for a lot of real work: client data can't go to someone else's cloud. NDAs, air-gapped environments, regulated data, or plain lack of trust. "Just use ChatGPT" isn't an option.

I built Nexus, an autonomous red+blue agent that addresses both. Here's how.

The idea: an evidence gate

An LLM is great at some things and terrible at others, so the two jobs are split: