For years, the race to build artificial general intelligence has been dominated by one modality: language. GPT, Claude, Gemini. All of them learned to reason by digesting oceans of text. A new white paper from researchers at Google DeepMind, Harvard, and other leading institutions argues that approach might be incomplete, and that the path to AGI could run through what machines see rather than what they read.

The paper, titled “Visual General Intelligence: A White Paper” and published on arXiv as 2608.25924, lays out a research agenda for what the authors call visual general intelligence, or VGI. The core thesis: AI systems should learn directly from images, videos, and geometric data to understand, predict, and act in the physical world.

What the paper actually says

This is not a product announcement or a benchmark-beating model reveal. It is a position paper, a collective argument from more than 21 researchers about where the field should invest its attention next.

Among the contributors are Robert Geirhos from Google DeepMind and Yilun Du from Harvard. The work grew out of discussions at the CVPR 2026 Visual General Intelligence Workshop, one of the premier gatherings in the computer vision community.