I gave two AI models the same 200 pieces of code, the same prompt, the same

question. One of them removed 51% of the false alarms. The other removed only

20% of the false alarms — and confirmed 90% of everything it was shown.

Same inputs. Same instructions. A 2.5× difference in the only thing I was

measuring.