I handed an AI an analysis that only counts if all seven of its steps run, in order. The procedure is written down. Here is what came back:

This method shows no measurable effect.

Fine, I thought. Dead end. Then, out of habit, I went and looked at what the AI had actually done.

It hadn't run it.

Out of everything in scope, the full seven-step procedure had executed exactly once. The rest was handled by a shortcut built from four numbers pulled out of a different source. That shortcut's result came back to me labeled as the conclusion of the seven-step method.