Instead of replacing the main model with a more expensive multimodal model, I gave a text-only agent an on-demand visual sensor.
Many coding agents can already read repositories, write code, execute commands, and run tests. But in real development workflows, they often get stuck on a surprisingly simple problem: the key information exists only in a screenshot.
For example:
Terminal error screenshots
Broken page layouts






