This is a cross-post. The original (with diagrams, code, and cheat sheets for every chapter) lives on my blog:

Vision-language models can describe a warehouse photo in fluent prose — and then fail to answer "how many meters is the forklift from the shelf?" They see but don't perceive. Over 2024–2026 a line of research closed that gap step by step, and the arc is remarkably coherent once you read it through a single lens:

The representational mismatch between discrete language tokens and continuous 3D geometry — and the successive attempts to close it.

I wrote a 10-chapter course tracing that arc paper by paper. Here's the map, with a one-line "why this paper had to exist" for each. Full chapter (analogies + mermaid diagrams + runnable code) is linked on each.

The 10-paper arc