An AI search answer can be fluent, current-looking, and wrong in a way that takes an hour to discover.

The dangerous part is usually not a completely invented citation. It is a real source attached to a sentence the source does not actually support. A press release becomes evidence for a market-size claim. A search snippet becomes evidence for a pricing detail. A paper's abstract becomes evidence for a conclusion that appears only after reading the methods section.

That is why I no longer evaluate AI search tools by asking which one writes the nicest answer. I evaluate the handoff from answer to evidence.

The Citation Test

Give every candidate the same set of real questions, then inspect the citations instead of the prose. For each claim that matters, record four outcomes: