Google DeepMind has released Gemini Robotics 2, a new version of its Gemini artificial intelligence model designed to control various robots, including humanoid machines [claim:1].

The system combines several different AI models to allow robots to understand and act within their surroundings [claim:2]. Gemini Robotics 2 utilizes a vision language model (VLM) for processing images and video, alongside two vision language action (VLA) models to control movement [claim:3].

To train the model, Google DeepMind used a combination of video examples, simulations, and human teleoperation [claim:4].

Embodied Reasoning and Capabilities

Gemini Robotics ER 2 functions as an embodied reasoning model that serves as a high-level brain for robots [claim:7]. This allows machines to plan multi-step tasks, understand the physical world, and engage in chat with humans [claim:7]. The model also enables multi-robot collaboration, allowing different types of robots to work together in shared spaces [claim:10].

The model reportedly can natively call tools, such as Google Search or other user-defined functions [claim:8]. It also integrates into the Gemini Live API through a bidirectional streaming endpoint for tasks sensitive to latency [claim:12].

Demonstrations and Hardware Integration

In a demonstration, Apptronik’s Apollo 2 robot used hands from Sharpa to tidy shelves [claim:5]. The Gemini Robotics 2 model is capable of controlling the SharpaWave hand—a five-fingered hand with 22 degrees of freedom—on the Apollo 2 robot for delicate tasks such as tying knots [claim:6].

Google DeepMind also demonstrated the model orchestrating Spot APIs from partner Boston Dynamics to create an interactive robot capable of fetching objects [claim:13].

The model can adapt to new bi-arm robot embodiments, typically requiring fewer than 200 examples and a few hours of adaptation time [claim:14]. Additionally, Gemini Robotics On-Device 2 provides a VLA model optimized to run locally on robotic devices without internet connectivity or network latency [claim:20].

Safety and Availability

Google DeepMind is introducing ASIMOV-Agentic, a new benchmark to measure the safety of AI systems that collaborate to control a robot [claim:15]. Gemini Robotics ER 2 is described as the company's safest robotics model to date regarding human proximity benchmarks and following safety constraints [claim:16].

Gemini Robotics ER 2 is available to developers through Google AI Studio, the Gemini API, and in private preview on the Gemini Enterprise Agent Platform [claim:9].

The model reportedly utilizes progress classification to track task completion across five levels: 0-20%, 20-40%, 40-60%, 60-80%, and 80-100% [claim:11].

Demis Hassabis, CEO of Google DeepMind, expressed hope to develop an AI operating system for various robots, similar to how Android operates smartphones [claim:19].