Eleven hours after the model-comparison post went up, a reader named Vinh Nguyen left a comment that unraveled its headline:
FATAL 0 against FATAL 0 is a tie in counts, but it is 29 items run once each. [...] Zero fatal out of 29 single trials is consistent with a true fatal rate up to about 10% at 95% confidence. [...] Running the whole exam five times per model and comparing fatal rates per item would tell you whether the tie is real or just the resolution of one pass, and it is the same order of cost as the run you already did.
He's pointing at a self-contradiction I published without noticing. My own post says, in the temperature section: "final scores should come from multiple runs, reported as a pass rate." And then the comparison itself — the whole "won by exactly one question" headline — sat on one pass per model.
No defense available. So I ran it.
The rerun






