Sooner or later, someone across the table asks the question. A client's security team. An auditor. An acquirer's diligence engineer. A regulator. Or you, at two in the morning:

"Why is this agent allowed to do that?"

Not does it work — demos answer that, and demos are cheap. Not how good is it — benchmarks answer that, about someone else's test set. The question is why it is allowed: what, concretely, stands behind the decision to let a language model operate near your customers and your money, and would that reasoning survive being read back to you a year later.

We build AI agents that transact — chat, voice, phone, SMS, one deterministic core underneath all of it. In an earlier essay we argued that the language model should conduct the conversation while that core makes every decision. In a second, that the model which writes must never be the model that grades. This essay is the third act, and it asks what the first two were building toward: when an agent is done, what actually admits it to production?

Our answer is unglamorous. Not a launch review, not a leaderboard, not a sign-off meeting. A file.