The setup was textbook. Four LLM judges on different base models. Double-blind pairs. Both presentation orders, to cancel position bias. Gold anchors seeded into the pool — samples where humans had already reached a verdict, including character-card copy a real user had flagged as "obviously AI" on sight, and community cards written by humans that had accumulated over a million conversations.
All the judges had to do was answer, for the anchors, "which of these two reads more human?" They scored 12%.
Failure one: not noisy — inverted
Twelve percent is the interesting kind of failure. A coin flip scores 50%; a noisy judge hovers around it. Scoring 12% means the panel was confidently wrong in a consistent direction: the copy humans called AI, judges praised as "grounded, textured, clearly human"; the human-written hits, judges dismissed as "formulaic, likely generated."
Failure two: agreement without accuracy







