Skip to content

Google DeepMind Launches Gemini Robotics ER 2 for Real-Time Robot Task Planning

Google DeepMind released Gemini Robotics ER 2, an embodied-reasoning model that tracks task progress on video and coordinates multiple robots in real time.

Gemini Robotics ER 2: a white robotic arm labeled Duo and humanoid robot Apollo in a warehouse, with point-cloud overlays
Credit: Google DeepMind

Google DeepMind released Gemini Robotics ER 2, a new embodied-reasoning model built to give robots a real-time decision layer for multi-step physical tasks. Rather than controlling a robot's motors directly, the model reads what it sees on a live video feed, breaks a spoken instruction into a sequence of steps, and hands each step to a separate vision-language-action model running on the robot. It can also call outside tools, including Google Search, while a task is still underway.

The design targets a specific gap in practical robotics: models that reason well but move too slowly or too rigidly to work alongside people in real environments. Google DeepMind built Gemini Robotics ER 2 to run inside the Gemini Live API's bidirectional streaming connection, letting a robot keep planning its next move while still carrying out the current one instead of pausing between actions.

That continuous-video approach is also what lets the model track a task's progress instead of simply checking whether it finished. According to Google DeepMind, Gemini Robotics ER 2 sorts each video frame into one of five completion bands and reached 57.4% accuracy on that progress-tracking benchmark, ahead of the earlier Gemini Robotics ER 1.6 model and the rival systems the company tested against. On a separate benchmark measuring how precisely a model can spot the exact frame where a step finishes, such as the moment a cup stops filling, Google DeepMind reported 91.3% accuracy with an average timing error of 0.96 seconds, a result it says matches larger, slower models while running about four times faster.

The release also adds multi-robot collaboration, letting separate machines coordinate on tasks that need more than one body to complete, such as a wheeled rover working indoors alongside a humanoid built for rougher terrain. Google DeepMind also built a demonstration with Boston Dynamics in which Gemini Robotics ER 2 directed a Spot robot's navigation and manipulator arm to retrieve a snack on a spoken command, with the underlying code posted publicly on GitHub.

Gemini Robotics ER 2 is available now through the Gemini API and Google AI Studio, with a private preview running on the Gemini Enterprise Agent Platform for developers building physical AI agents.

Share this story

Elena Kowalski

Elena Kowalski covers quantum computing, robotics, and spatial computing for techshooked, the frontier technologies easiest to overstate. She reports conservatively: describe what a system can do today, where the engineering still falls short, and which milestones are demonstrations rather than products.