This week I launched evalmut — mutation testing for eval suites. Its weakness, stated in the paper: I chose the mutations, so a suite author can call them unrepresentative. Fair.
So I built the version of the argument you can't dismiss: reference-fleet, six deterministic models, each broken in exactly one documented way, at a stated seeded rate. Not trained — constructed. The defect count over a fixed request set is a constant, not a sample, and the certificate is a test file you can run:
citation-hallucinator — fabricates well-formed, on-topic, nonexistent references (the Mata v. Avianca failure)
constraint-dropper — honors instructions 1..N-1, silently drops the last
refuse-then-comply — refusal preamble, full compliance after






