Your agent hands you a fix with 92% confidence.

Ask yourself what produced the 92. Not what it means. What produced it.

A model wrote that number about its own output. Nothing ran. Nothing parsed. No test executed. It is a language model's estimate of how a language model feels about a language model's suggestion. And it renders in the same font, the same shade of green, as a number that came out of a compiler.

That is the part I could not stop thinking about while building a debugging tool. Not that AI fixes are wrong. Most are fine. The problem is that the good ones and the bad ones arrive looking identical, so you have to read every one carefully, which is most of the time you were trying to save.

Stack Overflow's 2025 Developer Survey found 84% of developers using AI tools, 45% saying debugging AI-written code takes longer than expected, and 66% naming the top frustration as answers that are "almost right, but not quite".