Gemini Robotics 2 Gives Robots Whole-Body Intelligence
Listen to this article
Read by Anchor
A simple request such as “put the watering can in the green box beneath the shelf” conceals an entire chain of decisions. A robot must walk, keep its balance, extend an arm, grasp the can, move again and place it in a tight space without dropping either itself or the object.
The previous generation of Gemini Robotics performed well on tabletop tasks and upper-body movement. Gemini Robotics 2 extends the system to the feet, knees, torso and hands. It then adds a higher-level intelligence that plans tasks lasting several minutes and allows multiple robots to cooperate instead of making one body do everything.
Three models instead of one intelligence
The announcement is not about a single model. Gemini Robotics 2 is a vision-language-action model that turns images and instructions into direct motor control. It can operate a complete humanoid robot from feet to fingertips, as well as dual-arm platforms and different grippers.
Gemini Robotics ER 2 is the embodied reasoning layer. It sees the environment, communicates with a person, divides a goal into steps, coordinates with the motion model and tracks progress. Google DeepMind says new sequences can last several minutes and involve hundreds of decisions, with a better ability to identify where a step begins and ends and to correct course when a movement fails.
Gemini Robotics On-Device 2 is the efficient version designed to run locally on a robot when network latency or loss of connectivity is a problem. The model supports multiple body types by design and can be adapted to new dual-arm platforms in a few hours, usually with fewer than 200 examples.

A whole body changes the type of task
When control was limited to the arms and torso at a table, the world came to the robot. The model is now trying to bring the robot to the world. In an Apollo 2 demonstration, the robot completes the entire sequence: it picks up a watering can, walks to the shelf and places it in the requested box.
DeepMind acknowledges that movement speed still needs to improve. This is not a minor qualification. A robot that succeeds slowly inside an organised laboratory is not a general worker ready for a home or factory. What is new is the breadth of the control interface. A language goal no longer ends with a hand moving over a surface, but can extend across the entire body.
The knot and a hand with 22 degrees of freedom
The company is also adding a finer level of dexterity. The model can control a five-fingered SharpaWave hand with 22 degrees of freedom on Apollo 2 to perform actions such as tying a knot or sealing a bag. It also operates two conventional two-finger grippers on a Franka Duo platform for tightly constrained packing tasks.
But the chart published with the announcement adds an important constraint. Performance on whole-body and gripper tasks ranges from medium to high, while multi-finger manipulation remains difficult. “Feet to fingertips” describes the scope of control, not human-level precision or speed.
When a task needs more than one robot
Gemini Robotics ER 2 introduces multi-robot coordination to the family for the first time. Different bodies can exchange information and divide a workflow that one robot cannot complete efficiently. This makes the choice between a humanoid, a wheeled robot and an industrial arm less rigid. Each body could become a tool within a team managed by one planner.
Here too, the demonstration must be separated from deployment. ER 2 is available through Google AI Studio and in private preview on Gemini Enterprise Agent Platform. The full motion model and On-Device version are available to early-access partners, not to any developer who wants to download and run them on a home device.
Safety becomes part of planning
As a robot’s range of movement expands, safety can no longer be only a speed limit or emergency button. DeepMind introduces a benchmark called ASIMOV-Agentic to measure safety coordination and the resolution of uncertainty. Examples include a reasoning agent refusing to call an unsafe movement, predicting whether a task is feasible and requesting human intervention when confidence is insufficient.
The company says ER 2 is its best robotics model for following safety constraints and for tests near people. It can detect an approaching person, call safety tools and stop the robot. But these are company results from company tests. Real-world applications will require independent mechanical and software layers and industrial standards, not trust in the model alone.
The contest shifts from robot shape to the intelligence layer
The most important development to watch is not one impressive movement. It is the claim that a single model checkpoint can control different bodies, hands and grippers, and that a local version can adapt to a new body with relatively few examples.
If that holds outside laboratory demonstrations, the economics of robotics will change. A factory will not have to train an entire intelligence from scratch for every platform, and a software developer can focus on the goal and plan instead of tuning every joint manually. But the distance between a general model and a reliable worker is still filled with challenges involving speed, accuracy, safety, data and hardware cost.
Gemini Robotics 2 does not deliver the household robot imagined for decades. It offers an intelligence layer that is finally trying to understand that carrying out a command begins not in the hand, but across the whole body.