I publish free leaderboards of which products AI answer engines name when someone asks them to recommend software in a category. The method is deliberately boring: take one buying question, write it 44 different ways, ask every engine all 44, count every product named across the answers, publish the counts and the raw runs.

Last week I posted a result that I could not explain: two boards over what I had assumed was one market — small-business CRM and open-source CRM — came back with nothing in common in their top 20s. Not reordered. Zero shared products.

The obvious objection, and the one I got, is that this says more about my engines than about the question. LLM output is noisy. Maybe I had measured two engines having a bad day.

So I ran the controls. There are two knobs — the engine and the wording — and you can hold each one still and turn the other.

The setup