I used to compare image models in the laziest possible way: give them the same prompt, put the outputs next to each other, and pick the one I liked most.
It felt reasonable. It also gave me bad conclusions.
The problem showed up when I tried to use those “winning” models for actual creative work. A model that looked fantastic on a cinematic prompt could be irritatingly unreliable on an ad. Another would produce a less impressive first image but follow instructions more closely and save me two or three retries.
That made me stop asking which model makes the prettiest image.
Now I care more about a narrower question:






