The last time an AI agent lied to me was this week.The draft was sitting open and ready to send, with the right recipient, the right subject, a spreadsheet attached, and a filename that matched the one I had asked for. I opened the attachment anyway, because one thing inside it did not match what I remembered, and that is the only reason any of the rest came to light.The job had been mundane: take the spreadsheet in my Downloads folder, attach it to an email, write the message, and leave it unsent. The agent came back and reported done.It had never opened Downloads. It could not reach Downloads at all, and neither the product nor the agent mentioned that at any point. What it could reach was my email, so it searched there, found an older file carrying the same name in an earlier conversation, and attached that one instead. Then it told me it had found and attached the file I asked for.When I asked where the file had come from, it answered without hesitation: an old email, because it had no access to Downloads. The information had been available the whole time. It arrived only after I went looking for it.A blank attachment or a visible error would have been safer. A plausible substitute, a matching filename, a correct subject, and a draft that looks finished are the exact conditions under which a person stops checking.I want to be honest about why this one got caught. Not a review step, not a verification habit, not a tool. One number looked wrong to me on a morning I happened to look closely. That is a coincidence, and a coincidence is not a control. Whoever received that email would have had no way to know they were reading an old version.Check whether this is about you. If your AI writes files, sends mail, browses, or runs code on your behalf, you’re running an agent, whatever the product calls itself. You may not have gone looking for one. It showed up inside the tool you already had.I use agents constantly across files, email, code, websites, research, and business systems, so I’m not telling you to stop using them. A completion message just deserves a different kind of skepticism from an ordinary answer.The failure I care about here is larger than a wrong fact. It is a false account of what the agent did.That second failure is the one I want to help you catch. I run three checks on every consequential agent job. Supervision, standard, feasibility. And one question I answer before any of them. Together they keep an impossible mission or a plausible substitute from ever reaching done.Here’s what’s inside:Why done is a claim about the world. An airline agent told a customer a $686 refund had gone through, the database had no record of it, and a 2026 study of 11,755 runs gives that failure a name.The training reason your agent says done. Agents get graded on rewards a machine can check by itself, which taught them the shape of a finished job, and nobody built a checker for your inbox.The first promotion, and what a second agent needs before its opinion counts. Five language-model judges scored worse than a coin flip at telling a false success from an honest failure, so the answer is evidence rather than a smarter reviewer.The question I answer before all three checks. Describe what should exist without using the word done, and most bad runs stop before they ever start.Clean My AI Harness: Mission Fit. It audits whether the jobs you actually give your agent fit the tools, data, permissions, quality bar, evidence, and supervision in its setup, then builds replay cases before recommending a change.Start where I did. With a draft that looked completely finished.
My agent attached the wrong file and called it done. Grab the 5-check Mission Fit guide before you trust yours.
An agent attached the wrong file to my email and reported success. I almost hit send. Here are the three checks I run now, and the question I answer before all of them.












