The first experiment measured whether Claude Code knew the right answer. This one measured whether the work actually succeeded.

In my previous experiment, I tested whether a memory layer could stop Claude Code from confidently repeating project-specific mistakes.

It did.

But there was an obvious limitation.

The tasks measured whether the answer was correct. They did not ask the agent to modify a repository, run the result, and pass a deterministic test.