Google DeepMind is positioning Gemini Robotics 2 as its most advanced Vision-Language-Action, or VLA, model for robots. The central claim is not that one robot form factor has won. Instead, the model is designed to support a range of embodiments, from bi-arm platforms to full humanoids, while enabling whole-body control, dexterity and coordination among multiple robots operating in shared spaces.
That cross-embodiment focus gives useful context to DeepMind’s behind-the-scenes discussion of a familiar robotics question: when should a machine use humanoid legs, and when might a wheeled design be the better fit? The answer, at least in the materials available, is not a product verdict on legs versus wheels. It is a signal that robot intelligence and robot hardware must be designed together. Gemini Robotics 2 is presented as software intended to work across different physical forms rather than being confined to a single humanoid body.
Google DeepMind’s official Gemini Robotics page describes the system as a VLA model that can translate visual and language inputs into robotic actions. The accompanying video, Gemini Robotics 2 brings whole body intelligence to robots, demonstrates dexterity and multi-robot collaboration, reinforcing the company’s focus on coordinated physical behavior rather than isolated demonstrations.











