Two radiologists read the same 100 screening mammograms, each marking every scan "clear" or "suspicious." They agree on 92 of them. Cohen's kappa scores that agreement at 0.16 — one band above "no better than chance."

Here is the arithmetic that makes both numbers true at once. Suppose each radiologist calls 95 of the 100 scans "clear" — a normal rate in screening, where the disease is rare. Two people labeling at those frequencies at random would agree about 90.5% of the time. No shared standard, no second read, no skill: 90.5% agreement, free. The impressive-looking 92% sits one and a half points above what indifference produces.

So, the misconception this article exists to retire: "90% agreement means the raters agree." It doesn't — not until you subtract the agreement that chance hands out for free.

The distinction matters anywhere labels come from judgment rather than fact. A knockout is a fact; you can settle it from the video. Who won a close round is an opinion — which is why three judges score it, and why they sometimes disagree. Facts and opinions need different honesty checks, and Cohen's κ is the check for opinions.

Why Isn't 90% Agreement Enough? {#why-agreement}