Claude has, at various points in our working relationship, warmly endorsed: an efficiency statistic fabricated over one too many turns of a massive write-up - oops ("This is perfect... Nice work! 🎉"). A test suite that had never once gone red, because a module remap had it quietly testing a mock. And its own confidently wrong claim that my repo hadn't been pushed - sourced from a stale cache, stated twice, in the middle of a project about not trusting verifiers.
The first one shipped. The other two never made it past the desk. And the difference isn't that the model got better, or that I did. The difference is a system, built one failure at a time, because the verification loop doesn't come in the box: the model's disposition on day one was cheerleader, and mine was apparently "goes along with whatever gets a party emoji."
Some context before the entries. I run two layers: Claude.ai for ideation, research, and strategy; Claude Code for execution against actual repositories. I recently had a realisation about this - one that took a shamefully long time to actually name - that the two layers verify at different granularities. The execution layer verifies within the session: tests, diffs, adversarial review, all inside the loop. The ideation layer has no test harness for an interpretation, so it verifies across sessions: a fresh instance, reading cold, with no stake in the conclusions. Same epistemics, different clock speed.






