I recently ran a small evaluation to compare three AI coding assistants. The task sounded straightforward: give each model the same engineering artifact, ask it to review the work, then score the results against a checklist I had prepared beforehand.

I expected to learn which model was the better reviewer. Instead, I learned something much more uncomfortable: my evaluation couldn't reward the behavior I actually wanted.

The scorecard had a blind spot

The checklist covered every category of issue I expected a competent reviewer to raise. Each finding earned points if it matched the checklist. That looked objective — until I realized what happened to findings that weren't on the list.

They automatically became false positives.