Most AI roadmaps rest on a quiet bet: that the labs will eventually ship a model safe and aligned enough to simply trust in production. Better training, better guardrails, one more version, and the thing behaves.
Here is the flaw in that bet. Even a perfectly aligned model cannot tell you who used it, on what data, under whose policy, or hand you a record you could show an auditor. Those are not facts about how the model behaves.
They are facts about how it was deployed, and they live entirely outside the weights. A better model answers a different question than the one regulators, auditors, and security teams are actually asking, and the gap between those two questions is where the real risk now sits.
Security engineering named this trap fifty years ago. A 1972 US Air Force study defined the reference monitor, the component that decides whether an action is allowed, and set three conditions for trusting it: it must be tamperproof, invoked on every access, and small enough to be completely verified.
The 💜 of EU techThe latest rumblings from the EU tech scene, a story from our wise ol' founder Boris, and some questionable AI art. It's free, every week, in your inbox. Sign up now!A modern frontier model is none of the three. Alignment tries to make the model enforce the rules meant to constrain it, from the inside, and a system cannot be the thing that governs itself. Governance has to sit around the model, not in it.









