I gave two AI models the same 200 pieces of code, the same prompt, the same
question. One of them removed 51% of the false alarms. The other removed only
20% of the false alarms — and confirmed 90% of everything it was shown.
Same inputs. Same instructions. A 2.5× difference in the only thing I was
measuring.






