Google DeepMind launches Gemini Robotics 2 for robot control
Google DeepMind says Gemini Robotics 2 can control arms, humanoids and multi-robot systems, with early access handled by waitlist.
By Renata Fuchs · Policy Reporter
· 3 min read
Google DeepMind has introduced Gemini Robotics 2, a new vision-language-action model intended to control robots ranging from tabletop arms to full-body humanoids. The company says the model is its most advanced robotics system of this type, positioning it as a control layer for machines that need to interpret visual input, understand instructions and take physical actions.
The release matters for robotics developers because Google is packaging more of its Gemini work around embodied systems rather than limiting the brand to text, image and software tasks. DeepMind said developers can request early access through a waitlist, so this is not a broad self-serve launch for the main Gemini Robotics 2 model.
What is Gemini Robotics 2?
Gemini Robotics 2 is a vision-language-action, or VLA, model. In practical terms, that means it combines visual perception, language understanding and action generation so a robot can respond to what it sees and what it is asked to do.
DeepMind describes the model as an “intelligence layer” for more adaptable robots. According to the company, Gemini Robotics 2 can support whole-body motion, fine motor control and coordination across more than one robot. Those claims are broad, and DeepMind did not frame the announcement around a single commercial robot, customer deployment or named hardware partner.
The range of supported form factors is the key point in the announcement. DeepMind says the model can be used with small tabletop robot arms as well as humanoid systems, which would put the same model family across very different mechanical constraints. A tabletop arm and a full-body humanoid have different control problems, so the company’s claim is less about one task and more about model generality across robot bodies.
What is Gemini Robotics ER 2?
Google DeepMind also introduced Gemini Robotics ER 2, a separate model focused on what the company calls embodied reasoning. DeepMind uses that term for a system’s ability to understand the physical world and decide what actions are appropriate based on that understanding.
ER 2 is meant to operate at a higher level than direct robot motion control. DeepMind says it functions as a planning and reasoning layer for robots, replacing Gemini Robotics ER 1.6, which the company released in April. Unlike the main Gemini Robotics 2 model, ER 2 is available through Google AI Studio.
The split between Gemini Robotics 2 and Gemini Robotics ER 2 suggests Google is separating lower-level action control from higher-level physical reasoning, at least in how it is presenting the product line. For robotics teams, that distinction matters because perception, planning and actuation often move at different speeds and have different safety requirements.
DeepMind has not disclosed pricing, customer names or a general availability timeline for Gemini Robotics 2 in the details it announced. For now, the concrete developer path is early-access applications for Gemini Robotics 2 and access to the ER 2 preview through Google AI Studio.
This story draws on original reporting from The Decoder.