Last week, I ran an experiment that failed.

The hypothesis was simple: syllogistic prompts ("Major premise → Minor premise → Therefore...") should make AI models internalize rules more deeply than imperative prompts ("You MUST..."). I designed 8 probes, ran them across 3 conditions, and...

Cohen's d = −0.148. Direction: ~50%. Bayes Factor: < 1 (supporting the null hypothesis).

Zero effect. Nothing. I was ready to scrap the whole idea.

Then three experts looked at my data and said the same thing: "Your measurement tool is broken."