Google DeepMind demonstrates Gemini Robotics 2 humanoid control

Google DeepMind demonstrates Gemini Robotics 2 humanoid control

Google DeepMind has introduced Gemini Robotics 2, a robotics model suite it says can control full humanoid robots across walking, crouching, reaching and object manipulation, including demos on Apptronik’s Apollo 2 humanoid. According to the July 30 post on deepmind.google, the same Gemini Robotics 2 vision language action checkpoint was used across Apollo 2 configurations with SharpaWave and Inspire hands, plus a Franka Duo dual arm setup with a Robotiq gripper.

The release is centered on three models. Gemini Robotics 2 is the VLA model that converts vision and language inputs into motor control. Gemini Robotics ER 2 is the embodied reasoning model, a vision language model used for communication, task planning and coordination. Gemini Robotics On-Device 2 is the smaller local model, designed to run on robotic hardware without relying on a network connection.

For humanoid developers, the core claim is whole body control. Google DeepMind says earlier Gemini Robotics models controlled a humanoid upper body for tabletop work, while the new model can control the full Apollo 2 body. In one described task, Apollo receives an instruction to put a watering can into a green bin on a lower shelf, then walks to a table, picks up the can, steps to the shelf and places it in the target location. The company also acknowledges that movement speed still needs work.

Reported results show uneven dexterity

Google DeepMind’s own charts give a useful, if still company reported, view of where Gemini Robotics 2 humanoid control appears strongest. On Apollo with Inspire hands, reported general whole body manipulation accuracy was 68.4 percent for picking up from a table, 45.7 percent from the floor and 76.3 percent from a shelf.

The hand results are more mixed. On Apollo with SharpaWave hands, the post reports 36 percent accuracy for screwing a bulb, 92 percent for unscrewing a bulb, 44 percent for tying a trash bag, 32 percent for a dustpan task and 40 percent for a ziplock task. Google DeepMind says multi finger dexterous manipulation remains challenging, which is consistent with several of those figures sitting below half success.

The Franka Duo gripper results were stronger in the reported tests: 74.2 percent for general pick and place, 78.9 percent for diverse tool kitting and 89.6 percent for precise insertion tasks. For operators comparing end effectors, the release reads less like a claim that five finger hands are solved and more like a snapshot of a widening gap between reliable gripper work and still fragile anthropomorphic manipulation.

Reasoning, team tasks and local adaptation

Gemini Robotics ER 2 acts as the higher level controller in the system. Google DeepMind says it can observe a room, reason through steps, coordinate with the VLA model, track progress, correct itself when a step fails and run task sequences lasting several minutes with hundreds of decisions. The company also says the model can coordinate multiple robots, including different robot types, so they can communicate and work on a workflow together.

Access is split by model. Gemini Robotics ER 2 is available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform. The VLA and On-Device models are available to early access partners.

The local model is aimed at applications where latency or connectivity are constraints. Google DeepMind says Gemini Robotics On-Device 2 can adapt to new dual arm robot embodiments with a few hours of data, typically using fewer than 200 examples, including robots with different shapes, sensors and degrees of freedom. The post cites Dexmate, SO101 and Trossen platforms for those adaptation demonstrations.

Safety claims are part of the release

Google DeepMind is also introducing ASIMOV-Agentic, a benchmark for agentic safety orchestration and uncertainty resolution. The benchmark measures whether an embodied reasoning agent can refuse unsafe VLA tool calls, predict whether a task is possible and request human intervention when uncertain.

The company says Gemini Robotics ER 2 is its safest robotics model so far on safety constraint following and human proximity benchmarks. The post says the model can detect nearby humans, trigger safety tool calls and bring a robot to a safe stop when a person approaches too closely.

The main caveat is that the source is Google DeepMind’s own release. It provides task percentages and access status, but not independent test results, deployment counts or broad commercial availability for the VLA model beyond early access partners. The humanoid question now is whether the reported Apollo 2 whole body control can hold up outside curated demos, especially at speeds useful for real operations.

Source: deepmind.google

Similar Posts

New! 2026 Humanoid
Robot Market Report

198 pages of exclusive insight from global robotics experts — uncover funding trends, technology challenges, leading manufacturers, supply chain shifts, and surveys and forecasts on future humanoid applications.

Aaron Saunders
Featuring insights from Aaron Saunders, Former CTO of Boston Dynamics,
now Google DeepMind
Get the Report
New Report

The Humanoid Robot Supply Chain

Supplier Strategy and Market Positioning 2026–2027

Get the Report
New Report

Humanoid Foundation Models

The brains are being rebuilt

Get the Report