Buyers have started asking ChatGPT and Gemini which software to use instead of opening ten tabs. So over the last few weeks I ran a simple experiment 241 times: for each software category, I asked the assistants the questions a buyer actually types, wrote down every product they named, and wrote down every web page the answer was built from.
The result is 241 category measurements, 1,446 answers read, and 5,037 product mentions, all from real runs I can show you. Three findings surprised me enough to write this.
The method, so you can repeat it
For each category I asked six buyer-style questions: the plain "best X", plus "best free X", "most affordable X", "best X for small teams", "what X should I use", and "X recommendations". Each question got one live web search, put to ChatGPT and Gemini (and Perplexity on the earliest runs). Then I recorded, per answer, every product named and every source domain the assistant cited. No opinions, no model asked to rank anything. Whether a product counts is a literal check for its name in the answer text.
You can do this by hand for your own category in an afternoon. That is the whole point: it is countable, not a vibe.






