In the first post of this series I wrote about making a 29-question exam and getting it wrong five times myself — three times in the answer key, twice in the grader. Ever since, one question kept nagging me: if an AI wrote the exam, how many times would it get it wrong?

So I counted.

The experiment — I ordered 50 questions from an author AI

Hamel Husain's evals essay — the closest thing this field has to a textbook — describes how to scale up test questions: split the feature into scenarios, then mass-generate the input sentences with an AI. After putting an LLM in the grading seat, this was the next piece to test for real.

I gave the author AI (Sonnet 5) three things: