Every coding agent ships with some version of the same promise: it only touches what you allow it to touch. A scoped working directory. An approval dialog for risky commands. A list of paths it pretends not to see.
I stopped reading those as promises and started reading them as hypotheses. A hypothesis you can falsify in an hour, on your own hardware, before the agent ever gets near a repository that matters.
This post is the falsification kit: a honeypot repository, a shell-based audit script, and a six-run experiment battery. Nothing here requires trusting the vendor's description of its own sandbox.
Configured isn't the same as enforced
When people say an agent is "sandboxed," they usually mean one of three very different mechanisms, and they fail in very different ways:






