I had a simple-sounding question: does ChatGPT recommend this business?
You'd think you just ask it. Ask ChatGPT "best personal injury law firm in NYC", see if the business is named, record yes or no.
That works exactly once. Ask again an hour later and you might get a different answer. Not slightly different — potentially a completely different set of firms and a completely different set of cited sources.
Which means the naive version of this measurement is worthless. You're not measuring visibility, you're sampling a distribution once and calling it a fact.
This is the same problem anyone gets when they try to test an LLM-backed feature. Your normal testing instinct — same input, assert on output — just doesn't apply. So here's how I ended up designing around it, and the numbers that came out, which surprised me.






