We handed the same eight reply drafts, the same scoring rubric, and the same instructions to two independent AI reviewers. One returned revise 2 of 8. The other returned revise 7 of 8.

If your first instinct is "one of them is broken," it was ours too. It's also wrong, and the actual explanation reshaped how we write rubrics for any AI-on-AI review — code review, tone review, product QA, all of it.

Where 2-vs-7 actually came from

We diffed every disagreement. Almost none were about facts. They clustered into exactly three causes:

1. Unstated tolerance. Our rubric said replies should be "concise: 1–3 sentences." Reviewer A applied it as written and failed four drafts. Reviewer B reasoned that a five-sentence reply to a five-point technical comment is proportionate, treated the rule as being about lecturing, not arithmetic, and passed them — while explicitly flagging the deviation and inviting us to overrule. Neither misread the rubric. The rubric contained a number where it meant a judgment, and each reviewer resolved the ambiguity differently — which is precisely what an unstated tolerance forces reviewers to do.