I ran 110 vision-extraction jobs against a synthetic invoice. The low-resolution modes of the newest models never once returned the correct document — and the way they failed is worse than random noise.

All measurements in this post are as of July 10, 2026. Vision pipelines change; if you're reading this later, re-run the test before trusting the numbers.

The invoice that would have passed review

Here is a fragment of what GPT-5.6 (Sol, low-detail image mode) returned when I asked it to extract a synthetic invoice from a PNG:

{