The first experiment measured whether Claude Code knew the right answer. This one measured whether the work actually succeeded.
In my previous experiment, I tested whether a memory layer could stop Claude Code from confidently repeating project-specific mistakes.
It did.
But there was an obvious limitation.
The tasks measured whether the answer was correct. They did not ask the agent to modify a repository, run the result, and pass a deterministic test.







