Nine of the thirteen corrections our migration code had ever written in production were wrong. We found that out by stopping the bug fixes and counting.
Some context. We were migrating short-term leave data — sick days, compassionate leave, the odd half-day — from a legacy processor to a new integration service. Same source system upstream, same consumers downstream, new table in the middle. The kind of migration you scope at a week.
Two days of it went into fixing bugs in a mechanism that could never have worked. This is how we finally noticed. Part 2 is what we replaced it with.
All identifiers are anonymised and the systems described generically. The numbers are real.
The setup






