Google DeepMind has released Gemini Robotics 2, a new version of its Gemini artificial intelligence model designed to control various robots, including humanoid machines.

The system combines several different AI models to allow robots to understand and act within their surroundings. Gemini Robotics 2 utilizes a vision language model (VLM) for processing images and video, alongside two vision language action (VLA) models to control movement.

To train the model, Google DeepMind used a combination of video examples, simulations, and human teleoperation.

Embodied Reasoning and Capabilities

Gemini Robotics ER 2 functions as an embodied reasoning model that serves as a high-level brain for robots. This allows machines to plan multi-step tasks, understand the physical world, and engage in chat with humans. The model also enables multi-robot collaboration, allowing different types of robots to work together in shared spaces.

The model reportedly can natively call tools, such as Google Search or other user-defined functions. It also integrates into the Gemini Live API through a bidirectional streaming endpoint for tasks sensitive to latency.

Demonstrations and Hardware Integration

In a demonstration, Apptronik’s Apollo 2 robot used hands from Sharpa to tidy shelves. The Gemini Robotics 2 model is capable of controlling the SharpaWave hand—a five-fingered hand with 22 degrees of freedom—on the Apollo 2 robot for delicate tasks such as tying knots.

Google DeepMind also demonstrated the model orchestrating Spot APIs from partner Boston Dynamics to create an interactive robot capable of fetching objects.

The model can adapt to new bi-arm robot embodiments, typically requiring fewer than 200 examples and a few hours of adaptation time. Additionally, Gemini Robotics On-Device 2 provides a VLA model optimized to run locally on robotic devices without internet connectivity or network latency.

Safety and Availability

Google DeepMind is introducing ASIMOV-Agentic, a new benchmark to measure the safety of AI systems that collaborate to control a robot. Gemini Robotics ER 2 is described as the company's safest robotics model to date regarding human proximity benchmarks and following safety constraints.

Gemini Robotics ER 2 is available to developers through Google AI Studio, the Gemini API, and in private preview on the Gemini Enterprise Agent Platform.

The model reportedly utilizes progress classification to track task completion across five levels: 0-20%, 20-40%, 40-60%, 60-80%, and 80-100%.

Demis Hassabis, CEO of Google DeepMind, expressed hope to develop an AI operating system for various robots, similar to how Android operates smartphones.