Instead of replacing the main model with a more expensive multimodal model, I gave a text-only agent an on-demand visual sensor.

Many coding agents can already read repositories, write code, execute commands, and run tests. But in real development workflows, they often get stuck on a surprisingly simple problem: the key information exists only in a screenshot.

For example:

Terminal error screenshots

Broken page layouts