Originally published at tddbuddy.com.
Related reading: BDD Was a Coordination Tax. AI Just Repriced It is the prior post that this one extends; TDD Already Does BDD, Without the Gherkin and The Bar for TDD Just Moved name the test discipline this argument assumes.
An eval is a test. The industry just put a lab coat on it.
That sentence is the entire argument. Everything that follows is structural detail. The eval-tooling wave arrived with new vocabulary (golden datasets, judges, scorers, rubrics), new file formats, new dashboards, new job titles, and a confident announcement that "TDD does not work for AI" because there is no single correct output. Both moves are mistakes. The first is a vocabulary reinvention that hides what the discipline already knew how to do. The second is a strawman of TDD that good TDDers abandoned twenty years ago when they started asserting on properties and invariants instead of exact strings.
The cost of the lab coat is not aesthetic. It is structural. Treating evals as a separate discipline produces a separate specification language, owned by a separate role, maintained in a separate tool, drifting from the code it claims to describe. That sentence should sound familiar. It is the exact failure mode of feature files in the BDD era, restated with model-shaped nouns. The industry deleted Gherkin and rebuilt it, line by line, under a new name. The coordination tax that disciplined TDD finally killed is being charged again.






