Google DeepMind has released Gemini Robotics 2, a new version of its Gemini artificial intelligence model designed to control various robots, including humanoid machines.
The system combines several different AI models to allow robots to understand and act within their surroundings. Gemini Robotics 2 utilizes a vision language model (VLM) for processing images and video, alongside two vision language action (VLA) models to control movement.
To train the model, Google DeepMind used a combination of video examples, simulations, and human teleoperation.
Embodied Reasoning and Capabilities
Gemini Robotics ER 2 functions as an embodied reasoning model that serves as a high-level brain for robots. This allows machines to plan multi-step tasks, understand the physical world, and engage in chat with humans. The model also enables multi-robot collaboration, allowing different types of robots to work together in shared spaces.
The model reportedly can natively call tools, such as Google Search or other user-defined functions. It also integrates into the Gemini Live API through a bidirectional streaming endpoint for tasks sensitive to latency.
Demonstrations and Hardware Integration
In a demonstration, Apptronik’s Apollo 2 robot used hands from Sharpa to tidy shelves. The Gemini Robotics 2 model is capable of controlling the SharpaWave hand—a five-fingered hand with 22 degrees of freedom—on the Apollo 2 robot for delicate tasks such as tying knots.
Google DeepMind also demonstrated the model orchestrating Spot APIs from partner Boston Dynamics to create an interactive robot capable of fetching objects.
The model can adapt to new bi-arm robot embodiments, typically requiring fewer than 200 examples and a few hours of adaptation time. Additionally, Gemini Robotics On-Device 2 provides a VLA model optimized to run locally on robotic devices without internet connectivity or network latency.
Safety and Availability
Google DeepMind is introducing ASIMOV-Agentic, a new benchmark to measure the safety of AI systems that collaborate to control a robot. Gemini Robotics ER 2 is described as the company's safest robotics model to date regarding human proximity benchmarks and following safety constraints.
Gemini Robotics ER 2 is available to developers through Google AI Studio, the Gemini API, and in private preview on the Gemini Enterprise Agent Platform.
The model reportedly utilizes progress classification to track task completion across five levels: 0-20%, 20-40%, 40-60%, 60-80%, and 80-100%.
Demis Hassabis, CEO of Google DeepMind, expressed hope to develop an AI operating system for various robots, similar to how Android operates smartphones.
Updates
Gemini Robotics 2 features expanded capabilities including the ability to operate diverse platforms like Dexmate, SO101, and Trossen, as well as executing task sequences lasting several minutes involving hundreds of decisions. Developers can now utilize low-level control interfaces such as VLA models or navigation APIs, while the new model achieves improved performance in moment-finding tasks and includes safety features like pausing work when detecting nearby humans. Additionally, early-access partners can access VLA and on-device models, with specific demonstrations showing the system managing dexterous tasks like tight packing and collaborating with Boston Dynamics' Spot robot.
Gemini Robotics ER 2 allows robots to process tasks and plan future actions simultaneously while executing long sequences involving hundreds of decisions. The system introduces 'moment-finding' capabilities to identify critical video frames and includes safety features that enable autonomous work pauses when humans are detected nearby. Google has also partnered with Boston Dynamics to integrate this technology into legged robots, with demonstration code now accessible on GitHub. Developers can further utilize the model on platforms such as Dexmate and Franka Duo, while early-access partners gain entry to VLA and On-Device versions that support advanced motion transfer.
Gemini Robotics ER 2 enables robots to execute multi-step tasks lasting several minutes without stop-and-think pauses by simultaneously performing actions and reasoning. The system integrates with platforms such as Boston Dynamics' Spot, Apptronik's Apollo 2, and Franka Duo, allowing for complex maneuvers like dexterous packing and autonomous safety stops. Google also introduced a new safety benchmark to evaluate foundation models acting as VLA orchestrators and confirmed that, while VLA and On-Device versions are available to early-access partners, there are no immediate plans for a consumer-facing robot rollout.
Gemini Robotics ER 2 enables robots to execute complex, multi-step tasks lasting several minutes without 'stop-and-think' pauses by processing decisions in real-time. The system supports diverse hardware, including Boston Dynamics' Spot and Apptronik's Apollo 2, and features a new safety benchmark that mandates environment monitoring and human clarification. Additionally, Google confirmed that these capabilities are now available to early-access partners through VLA and On-Device models, while clarifying that no consumer-facing robots are planned for the near future.
Gemini Robotics 2 allows for real-time multimodal streaming of video, audio, or text, enabling robots to execute long task sequences lasting several minutes without 'stop-and-think' pauses. The model can be evaluated through simulation, real-world control, or remote human pairing, and it achieves significant gains in 'moment-finding' performance. Additionally, the system enables humanoid robots to perform actions like walking and crouching, supports various platforms such as Dexmate and SO101, and uses the ASIMOV-Agentic benchmark to ensure safety by predicting task feasibility and requesting human intervention when uncertain.
Gemini Robotics 2 enables robots to execute lengthy task sequences involving hundreds of decisions and allows for the integration of low-level control interfaces, such as navigation APIs, as tools. The model facilitates real-time multimodal streaming and can be evaluated through simulation, real-world control, or human-remote pairing. Additionally, the system supports various hardware, including the ability to operate standard two-fingered parallel grippers on a Franka Duo platform and adapt to diverse embodiments like Dexmate, SO101, and Trossen.
Gemini Robotics ER 2 enables robots to execute long task sequences lasting several minutes through multimodal inputs, including video, audio, or text. The model achieved 57.4% accuracy in task progress classification and 91.3% in moment-finding tests, with an average error margin of 0.96 seconds for event timing. Additionally, Google has released the ASIMOV-Agentic safety benchmark on Hugging Face to evaluate an agent's ability to monitor environments and request human intervention.
Gemini Robotics ER 2 can perform multi-step tasks lasting several minutes through a vision language model architecture that allows humanoid robots to execute actions like walking, crouching, and stretching. The model achieved 91.3% accuracy in moment-finding tests with an average timing prediction error of 0.96 seconds, and its new ASIMOV-Agentic safety benchmark is now available on Hugging Face. Additionally, while Google is not planning immediate consumer-facing releases, VLA and On-Device models are already available to early-access partners.
Gemini Robotics ER 2, its predecessor version 1.6, functions as a vision-language-action (VLA) model that enables robots to execute multi-minute task sequences and self-correct by readjusting motions. The system achieved 91.3% accuracy in moment-finding tests and 57.4% accuracy in task progress classification, while its new ASIMOV-Agentic safety benchmark is now available on Hugging Face. Additionally, the model's capabilities have been demonstrated on various platforms, including Apptronik's Apollo 2 robot and Boston Dynamics machines.
Gemini Robotics ER 2 is a significant upgrade over version 1.6 and functions as a vision-language-action (VLA) model capable of executing multi-step tasks lasting several minutes. The system features a multi-layered safety approach, including the new ASIMOV-Agentic benchmark available on Hugging Face, which evaluates an agent's ability to refuse unsafe calls and seek human intervention. In performance tests, the model achieved 91.3% accuracy in moment-finding and 57.4% accuracy in task progress classification.