Physical AI just crossed a threshold that matters. When Google DeepMind released Gemini Robotics-ER 1.6 in April 2026, the headline benchmark told the story clearly: 93% accuracy on industrial instrument reading — compared to 23% for the prior version and 72% for Gemini 3.0 Flash on the same task. Boston Dynamics has already deployed it on its Spot quadruped robot platform, live for all AIVI-Learning customers as of April 8, 2026. And unlike most frontier robotics research, this one comes with developer access: ER 1.6 is available via the Gemini API and Google AI Studio, with a public Colab notebook and configuration examples. This guide covers what changed, how the agentic vision architecture works, how Boston Dynamics is using it in practice, and how to start building physical AI applications yourself.

Background: The Gemini Robotics Model Family

Gemini Robotics is Google DeepMind's line of vision-language models designed for physical systems. The family splits into two branches with different purposes:

Gemini Robotics VLA (Vision-Language-Action): A generalist model that outputs physical actions directly, controlling robot actuators end-to-end. Designed for manipulation tasks and general-purpose robot control.