I used to compare image models in the laziest possible way: give them the same prompt, put the outputs next to each other, and pick the one I liked most.

It felt reasonable. It also gave me bad conclusions.

The problem showed up when I tried to use those “winning” models for actual creative work. A model that looked fantastic on a cinematic prompt could be irritatingly unreliable on an ad. Another would produce a less impressive first image but follow instructions more closely and save me two or three retries.

That made me stop asking which model makes the prettiest image.

Now I care more about a narrower question: