Gemini Robotics 2: one Google DeepMind brain for different robots
Google DeepMind presented Gemini Robotics 2 - a new family of AI models designed to control robots. The most important change is the extension of control from the arms and upper body to the entire humanoid: from the feet to the fingers. The company also shows longer task planning, cooperation of several robots and a model that works locally, without a permanent connection to the cloud.
Highlights
- Gemini Robotics 2 is designed to control a variety of structures: from tabletop robots and arms to full humanoids,
- on the Apptronik Apollo 2 humanoid model controls walking, bending, balancing, reaching and manipulating objects,
- Gemini Robotics ER 2 is responsible for understanding the environment, talking to humans and planning multi-stage tasks,
- the system can coordinate several robots performing a common task,
- Gemini Robotics On-Device 2 can run locally and, according to Google, adapt to the new robot design with less than 200 examples,
- ER 2 scheduling model is available in public preview, but direct traffic control models remain in Early Access,
- Google has not introduced its own robot for sale - it is developing an intelligence layer for partner equipment.
Three models instead of one product
The Gemini Robotics 2 name includes three related elements. The first is a VLA, or vision-language-action, model. It receives an image, instructions and data from the robot, and then generates movement commands. It is responsible for direct control of the arms, hands, grippers, legs and balance.
The second element, Gemini Robotics ER 2, acts as a master "brain". It analyzes the situation, breaks down the command into stages, tracks the execution of the task and decides what should happen next. It does not control each engine directly, but delegates tasks to a lower motor layer.
The third model, Gemini Robotics On-Device 2, has been optimized to run directly on robot hardware. It is intended to reduce delays, enable operation without the Internet and facilitate the implementation of the system on new structures.
Entire humanoid body under the control of one model
In previous versions of Gemini Robotics, demonstrations focused primarily on table manipulation. Now Google DeepMind shows Apptronik Apollo 2 performing tasks that require the simultaneous use of its legs, torso and arms. The robot can approach an object, bend down, maintain balance, move it and put it back in the indicated place.
In one demonstration, Apollo 2 is instructed to place a watering can in a green container on the lower shelf. The robot must locate the item, approach it, grab it, move to the shelf, lower its body position and perform the putting away.
Google also released results for general full body manipulation. In tests with Apollo and Inspire hands, the average effectiveness was 68.4%. when lifting objects from a table, 45.7 percent from the floor and 76.3 percent off the shelf. This is significant progress, but at the same time it is proof that the system is not yet reliable in every scenario.
Great hands, but difficult tasks still cause problems
Gemini Robotics 2 also controls a five-finger SharpaWave hand with 22 degrees of freedom. Demonstrations include tying a bag, closing a ziplock bag, and screwing in and unscrewing a light bulb.
The most valuable part of the publication, however, are not the videos, but the results provided by Google. Unscrewing the bulb reached 92 percent. effectiveness, but its screwing in only 36%. Tying the bag was successful in 44 percent of cases. trials, closing a ziplock bag in 40% and using a scoop in 32%.
This is a good indication of the current state of robotics. The model can perform tasks that were previously beyond the reach of general-purpose systems, but the most precise tasks are still significantly less reliable than human work. Results are from Google DeepMind testing and have not yet been independently confirmed in long-term deployments.
Minute planning and robot collaboration
Gemini Robotics ER 2 is designed to perform longer sequences involving hundreds of decisions. The system observes the surroundings, assesses progress and can repeat the step if the previous attempt fails. It also recognizes the start and end times of individual stages.
What's new is the cooperation of several robots. The models are intended to recognize the capabilities of individual structures and divide tasks. In practice, one robot could perform precise manipulation and the other could transport elements or work in a different area.
Such a model may be important for factories and warehouses because future automation will likely not be based on a single universal humanoid. A team of specialized devices using a common planning system is more realistic.
Local model: less latency and more control over your data
Gemini Robotics On-Device 2 is designed to run directly on the device. This is important in applications where an interruption in Internet access cannot stop the robot or transmitting images from cameras to the cloud would be too slow or problematic from a privacy point of view.
Google says the model can be adapted to the new two-arm robots in a matter of hours, typically using fewer than 200 examples. The company demonstrates operation on structures that differ in shape, sensors and the number of degrees of freedom.
This does not mean, however, that any manufacturer can download the model and run it on their robot today. On-Device 2 is only available to select testers, and the direct VLA model requires early access enrollment.
What is available and what remains a demonstration
Gemini Robotics ER 2 is now available in Google AI Studio and via Gemini API as a public preview. Companies can experiment with its ability to analyze image, video, audio and action planning.
The most spectacular elements - full humanoid control and a local motion model - are not yet a generally available product. Google works with hardware partners including Apptronik, Boston Dynamics and Agile Robots, and access to motion models remains limited.
No pricing, licensing terms for commercial fleets, or timeline for wide release provided. It is also unknown how much data and integration work will be needed to achieve high efficiency in a specific plant.
Security: new benchmark and human arrest
Google presented the ASIMOV-Agentic benchmark, which is intended to assess the system's ability to reject dangerous commands, recognize situations that cannot be performed safely and ask a human for help.
Gemini Robotics ER 2 is to better detect the presence of a human and trigger a safe stop of the robot. The company reports high results in its own tests of compliance with restrictions and reaction to a person located one meter away.
The AI model, however, does not replace classic safeguards: force limitation, speed control, emergency stop, safe motion planning and certification of a specific machine. Security depends on the entire system, not just the Gemini layer.
Why is it important for the robot market
Google DeepMind's greatest ambition is not to sell a single humanoid. The company is trying to create a layer of intelligence that can be transferred between different designs. If this model proves successful, equipment manufacturers will not have to build the entire language understanding, planning and behavior learning system from scratch.
This could accelerate market growth, just as common operating systems accelerated smartphone development. For today, however, this is a direction, not a ready standard. Models remain semi-closed, integrations require partnerships, and the reported results still show a significant difference between a successful demonstration and the reliability required in your home or business.
RoboMorrow assessment
Status: Major technology launch, limited commercial availability. Gemini Robotics 2 is an important step because it combines full-body control, precise manipulation, longer planning and the ability to act locally. The most valuable thing is that Google also published results showing the limitations of the system.
However, this is not a "ready-made brain" that can be installed in any humanoid today. Companies interested in implementation should observe the availability of the affiliate program, hardware requirements, licensing and independent results from real factories and warehouses.
Featured image from official Google DeepMind materials for Gemini Robotics 2.
Sources
Google DeepMind: Gemini Robotics 2 brings whole body intelligence to robots
Google DeepMind: Gemini Robotics 2 - models and capabilities
Google DeepMind: Gemini Robotics ER 2
Google DeepMind: Gemini Robotics On-Device 2 - model card