Google DeepMind Unveils Gemini Robotics 2 With Whole-Body Humanoid Control
Google DeepMind has released Gemini Robotics 2, a platform that extends its robotics AI from arm-and-hand manipulation to full-body humanoid control. The system combines a vision language model with two vision-language-action models and introduces a new safety benchmark called ASIMOV-Agentic.
Google DeepMind announced Gemini Robotics 2 on July 30, 2026, a platform that gives humanoid robots what the company calls "intelligent whole-body control" — a step up from the arm-and-hand-only manipulation of the original Gemini Robotics system released earlier.
Google released a video demonstrating robots performing tasks including inserting a cassette tape into a boombox, screwing in a lightbulb, tying up a garbage bag, cleaning up trash, and picking up watering cans. According to Google, the footage shows "real-time" performance by "fully autonomous" robots — a claim the company is making explicit contrast to competitors. Engadget notes this stands in contrast to Tesla's Optimus, where a high-profile demo was later revealed to have relied on teleoperators rather than autonomous control.
How it works
Gemini Robotics 2 runs multiple AI models simultaneously to give robots multimodal understanding of their environment, according to Google. The architecture includes one vision language model for scene understanding and two separate vision-language-action models — one governing full-body movement and another controlling hand movement.
These are not general-purpose robots. Each demonstrated task required specific training using a combination of human teleoperation, video examples, and simulation data. Google DeepMind has not disclosed parameter counts, training data volume, or a timeline for broader task generalization.
Safety layer: ASIMOV-Agentic
Google says it is introducing a new benchmark, ASIMOV-Agentic, designed to evaluate whether a given robot command could lead to a harmful outcome. The company describes a "multi-layered" safety approach with guardrails built into each model layer. Carolina Parada, head of robotics at Google DeepMind, told Wired that safety concerns intensify as robots are deployed in more varied, less-controlled environments: "There's a lot of uncertainty that will show up, and so you want to be able to understand the safety question more deeply."
No quantitative benchmark scores for ASIMOV-Agentic have been published.
Context: the "physical AGI" framing
Google DeepMind described the release as a step toward what it calls "physical AGI" — robots capable of performing any task a human can, according to a statement given to Wired. This is a company aspiration, not a near-term product claim. Google has not announced a consumer or commercial robot release tied to this platform.
The announcement arrives against a backdrop of aggressive claims from Tesla CEO Elon Musk, who has said Optimus will become "the biggest product ever," projecting $10 trillion in eventual sales. Musk previously predicted 50,000 to 100,000 Optimus units by 2026 and 1,000 units in manufacturing facilities by the end of 2025 — neither of which materialized, per Engadget's reporting.
What this means
Gemini Robotics 2 signals Google's shift from single-limb manipulation demos toward coordinated full-body humanoid control, using a layered multi-model architecture rather than one monolithic policy. The introduction of ASIMOV-Agentic suggests Google is treating safety evaluation as a first-class benchmark category for embodied AI, not an afterthought — notable given the physical stakes of deploying large robots near humans.
Still, every capability shown was task-specific and required substantial human-in-the-loop training data. That's a meaningful gap between this demo and general-purpose robotic autonomy. The real test will be whether Google publishes reproducible benchmark numbers, discloses model sizes, or ships hardware partners can actually deploy — none of which has happened yet. Until then, this is a research milestone with a marketing video attached, not a product.
Related Articles
AllSpark's Iris-mini and Iris-pro Top Open-Weight Search Agent Benchmarks
Chinese lab AllSpark has released Iris-mini and Iris-pro, two open-weight search agents built on Qwen3 models that claim the top spot among open-weight systems in their size classes on four research benchmarks. The release includes model weights, an agent harness, and evaluation code, with training pipelines to follow.
Tencent Open-Sources AuK, a 1.5B-Parameter Speech Generation and Editing Model
Tencent has open-sourced AuK, a 1.5B-parameter foundation model for speech generation and editing that handles TTS, content editing, and audio enhancement through natural-language instructions. The release includes a distilled AuK-Flash variant for 4-step fast inference, both under MIT license.
Google Releases TimesFM-3, a 330M-Parameter Model That Forecasts Sales Using Weather and Discount Data
Google Research has released TimesFM-3, a 330-million-parameter time series forecasting model that predicts outcomes like sales by combining related variables, historical data, and known future events such as discounts or weather. The model claims top rankings on three benchmarks against Amazon's Chronos-2 and the Toto-2.0 family.
DeepSeek Ships V4.1-Flash With Novel Encoder-Decoder Architecture, Cuts KV Cache to 1/8 of Predecessor
DeepSeek released V4.1-Flash, a 763B-parameter model built on a new causal encoder-decoder architecture that splits 8B active parameters for prefill and 16B for decode. The model adds native vision support, a 1M-token context window, and shrinks KV cache footprint to roughly 1/8 of DeepSeek V4 Flash, while retiring V4 Pro.
Comments
Loading...