Google has launched Gemini Robotics ER 2, an embodied reasoning model that helps robots understand their surroundings, communicate with people, and complete complex physical tasks in real time.
The model acts as a high level brain, planning tasks before passing physical execution to a vision language action model or another robot control system.
Gemini Robotics ER 2 can process continuous video, audio, and text while using tools such as Google Search, navigation systems, and developer defined functions. It can plan its next step while a robot is still acting, reducing pauses between decisions.
The model is available through the Gemini API and Google AI Studio, with a private preview offered through the Gemini Enterprise Agent Platform.
Gemini Robotics ER 2 can track task progress, identify mistakes, and adjust actions without restarting an entire workflow.













