I recently ran a small evaluation to compare three AI coding assistants. The task sounded straightforward: give each model the same engineering artifact, ask it to review the work, then score the results against a checklist I had prepared beforehand.
I expected to learn which model was the better reviewer. Instead, I learned something much more uncomfortable: my evaluation couldn't reward the behavior I actually wanted.
The scorecard had a blind spot
The checklist covered every category of issue I expected a competent reviewer to raise. Each finding earned points if it matched the checklist. That looked objective — until I realized what happened to findings that weren't on the list.
They automatically became false positives.






