The Test That Has to Fail

There is a specific kind of dread that sets in when everything looks too good.

I had spent weeks wiring an AI judge into the review layer of a generation pipeline.

The judge was an LLM we had carefully prompted to read outputs and score them against

a rubric: coherence, factual grounding, format compliance, tone. It was supposed to