I got a comment on the last post that I want to answer properly, because it gets at the real question. They agreed that memory adherence is a systems problem and not just a prompting problem, then asked whether I had tested the approach across model families (GPT, Claude, Gemini) and seen meaningful differences in write reliability.

There are differences, but they are not about the model. They are about how much of the turn each family lets you control. If adherence is a systems problem, then the thing that actually decides write reliability is which family hands you enough control surface to build the system in the first place.

Here is what I have found in practice. Treat this as a field report, not a benchmark, and an early one. I have not pushed every family equally hard, and that matters for how much weight the comparison can carry. More on that below.

Claude is where I have gone the deepest

This is the family I have spent the most time controlling, so weigh the claim with that in mind. What I can say is that Claude gives you a full ladder, from shallow to deep, and you can stop at whatever rung the job needs.