Google DeepMind has unveiled Gemini Robotics 2, a Vision-Language-Action (VLA) model designed to power full-body humanoid robots, extending beyond the upper-body functionality of the original Gemini Robotics system introduced last year1.
The model integrates visual, natural language, and physical action processing to enable robots to walk, crouch, reach, manipulate objects, and tidy rooms. Google said Gemini Robotics 2 can also coordinate with other robots to complete tasks faster.
In real-world testing applied to Apptronik's Apollo 2 robot, the system successfully executed commands such as moving a watering can from a bottom shelf into a green bin. Shelf-based object picking achieved 76.3% accuracy, while tabletop object picking reached 68.4% and floor-level picking came in at 45.7%.
Fine motor tasks using Scharff robotic fingers on the Apollo platform proved more difficult. Removing a lightbulb achieved a 92% success rate, but inserting a lightbulb dropped to 36%. Tying a trash bag succeeded 44% of the time, closing a zipper bag 40%, and using a dustpan 32%. Bloomberg separately reported on the dexterity challenges facing the Gemini AI robotics system2.
Alongside Gemini Robotics 2, Google introduced two additional models. Gemini Robotics ER2 is a reasoning model capable of predicting the start and end of tasks and planning multi-step operations. Gemini Robotics On-Device 2 is a lightweight variant designed to run without an internet connection.
ANALYSIS The benchmark results illustrate a clear gradient: gross motor tasks such as shelf picking are approaching reliable performance, while fine manipulation — inserting lightbulbs, tying bags, operating dustpans — remains substantially less reliable. The gap between coarse and fine dexterity defines the near-term engineering frontier for VLA-powered humanoids.
Market research firm KAI Research projects the VLA market will grow to $40.5 billion by 2035. Separately, the FCC announced a ban on future sales of Chinese-made robots due to security concerns.
ANALYSIS The introduction of an on-device model and a dedicated reasoning layer suggests Google is building a tiered architecture — cloud-connected for complex planning, edge-capable for latency-sensitive or offline environments. Combined with the full-body control upgrade, the Gemini Robotics 2 suite constitutes a bid to supply the foundational AI stack for general-purpose humanoid robots.