Google DeepMind introduced three specialized robotics models working in tandem: Gemini Robotics 2, a vision-language-action model that converts vision and language input into motor control for whole-body humanoid manipulation; Gemini Robotics ER 2, a reasoning layer that plans multi-step tasks lasting several minutes and coordinates multiple robots; and Gemini Robotics On-Device 2, which runs locally on robotic hardware and adapts to new robot designs with fewer than 200 examples.
