"+7 points from context" was an illusion — measure the judge's spread first

When I had an AI score quality, just measuring the same thing twice moved the number by 6.1 points. And that wobble was exactly the size of the "+7.2-point improvement" I was about to adopt. Put the two head-to-head and the win rate was 53% — indistinguishable from a dead heat.

A judge is an instrument, and every instrument has its own noise and its own bias. Don't make decisions off its output until you've measured that first.

In one line: measure the judge's spread first, and don't call any difference at or below that spread an improvement.

Turning quality into a number is hard. Whether a translation is good, whether a summary is on point, whether an answer is helpful — there's no yardstick for that kind of "good." So there's a move everyone reaches for now: have an AI do the scoring (LLM-as-judge; from here on, judge = the AI doing the scoring). Faster and cheaper than a human, and it will hand back "this one passes, this one fails" for 100 items or 1,000.