TL;DR. Almost every eval harness reports a pass rate with an error bar, and almost every one of those error bars comes from the normal approximation: p̂ plus or minus 1.96 times the square root of p̂(1 - p̂)/n. That formula is taught first, implemented everywhere, and reasonable near a pass rate of 50 percent. It falls apart at the extremes, which is precisely where any model worth shipping lives. At 49 of 50 passing it produces an upper bound of 1.0188, a probability above one. At 50 of 50 it produces the interval [1.0, 1.0], a claim of perfect certainty from fifty observations. Worse than either artifact: when the true pass rate is 98 percent and n is 50, its actual coverage is 63.5 percent. The reframing is that this is a solved problem, and has been since 1927. Invert the score test instead of approximating around the estimate, and you get the Wilson interval, which is one argument change in the library you already have installed.

The regression I shipped because I misread an interval

Three years ago I owned the eval suite for a document extraction pipeline. Fifty held-out documents, hand-labeled, each either extracted correctly or not. We were shipping a prompt change and the numbers looked clean: 49 of 50 passing, and the harness printed a 95 percent confidence interval of [0.941, 1.000].