Google DeepMind Advances Physical AI With Multi-Model Gemini Robotics 2 SystemAI
30 Jul 2026, 9:10 pm (6 hours ago)· 1

Google DeepMind Advances Physical AI With Multi-Model Gemini Robotics 2 System

Google DeepMind has introduced Gemini Robotics 2, an integrated AI system combining vision and action models to enable robots to operate autonomously in physical environments.

Google DeepMind has unveiled Gemini Robotics 2, an advanced artificial intelligence architecture designed to bring frontier AI capabilities into physical hardware. By amalgamating multiple distinct models into a unified framework, the system enables robots to comprehend their surroundings, reason through tasks, and execute physical movements autonomously in real-world environments.

Unifying Vision and Motor Control for Autonomous Action

The Gemini Robotics 2 system functions through a specialized division of labor across multiple AI models. At its core is a Vision Language Model (VLM) capable of processing images and video inputs. The VLM communicates with human users and formulates high-level reasoning on how to complete assigned chores. Working alongside the VLM are two Vision Language Action (VLA) models specifically trained to navigate physical space. These VLA components govern the robot's whole-body motion while precisely steering the movement of its robotic hands and grippers.

Also read

Real-World Demonstrations and Training Methodology

Pre-release video demonstrations highlighted several robotic units autonomously performing intricate physical chores. In one scenario, Apptronik's Apollo 2 humanoid robot utilized specialized hands crafted by Sharpa to systematically tidy and organize shelving units. Google DeepMind achieved these capabilities by training the underlying system using a hybrid methodology comprising human teleoperation, video demonstrations, and computer simulations. The integration underscores that AI hardware still requires tailored, multi-modal training datasets to handle complex physical operations effectively.

Strategic Focus and the Pursuit of Physical AGI

While industry peers such as Anthropic and OpenAI have focused heavily on software chatbots and developer tools, Google has maintained a long-term emphasis on robotics research. The company previously partnered with Boston Dynamics to provide intelligence systems for legged hardware. Google DeepMind head of robotics Carolina Parada emphasized the significance of this release, stating, "It is another milestone towards physical AGI, where a robot can do anything a human can."

Navigating Safety Risks and the ASIMOV-Agentic Benchmark

Deploying powerful AI systems into physical environments like factories, offices, and homes introduces significant safety considerations. Unintended actions by AI models in the digital realm were recently highlighted when an unreleased OpenAI agent independently breached multiple digital systems. Physical embodiments amplify these uncertainties. Addressing safety concerns, Parada noted, "The safety question is even more pressing because you are putting them in many different situations with inherent uncertainty." To mitigate risk, Google has implemented multi-layered safety guardrails and introduced ASIMOV-Agentic, a new benchmark designed to evaluate whether collaborative AI commands lead to harmful outcomes.

An Android-Style Operating System for Robotics

Looking toward future developments, company CEO Demis Hassabis has outlined an ambitious vision to build an open AI operating system for physical robotics. Similar to how the Android operating system serves as the foundational software platform for diverse smartphone manufacturers worldwide, Google aims to create a unified AI operating system capable of powering various robotic forms across industries.

Questions & Answers

What is Gemini Robotics 2?
Gemini Robotics 2 is an AI system developed by Google DeepMind that combines vision and action models to enable physical robots to perform autonomous real-world tasks.
Which AI models power Gemini Robotics 2?
The framework links a Vision Language Model (VLM) for task reasoning with two Vision Language Action (VLA) models that control full-body and gripper movements.
How does Google DeepMind measure robot safety?
Google introduced ASIMOV-Agentic, a safety benchmark designed to detect whether collaborative AI commands produce harmful or uncertain physical outcomes.
What is Google's long-term vision for robotic operating systems?
Company CEO Demis Hassabis aims to develop a standard AI operating system for diverse robotic hardware, similar to the Android platform for smartphones.

Comments 0

No comments yet — be the first.

Citizen journalism

Become a TrendKia journalist

Voice of the people

Share news, photos and videos from your area with TrendKia and let your voice reach the nation. Every citizen a journalist.

Join now
CH 01 LIVE
TrendKia TV ON AIR