AI2 releases robotics models trained entirely in simulation, achieving zero-shot real-world transfer
AI2 has released MolmoSpaces and MolmoBot, robotics models trained exclusively in simulation that transfer directly to real robots without manual real-world data collection or fine-tuning. The approach eliminates months of teleoperated demonstrations typically required for simulation-trained robots. Both systems are open-source.
AI2 Releases Robotics Models Trained Entirely in Simulation, Achieving Zero-Shot Real-World Transfer
AI research institute AI2 has released two open-source robotics models—MolmoSpaces and MolmoBot—trained exclusively in simulation that achieve zero-shot transfer to real robots without any manually collected real-world data or fine-tuning.
What's New
The models represent a significant shift in robotics training methodology. Conventional approaches require researchers to spend months collecting teleoperated real-world demonstrations before simulation-trained robots become reliable. AI2's approach eliminates this bottleneck entirely.
MolmoSpaces, the foundation dataset, contains:
- 230,000+ indoor scenes
- 130,000+ curated objects
- 42 million physics-based robotic grasping annotations
MolmoBot, built on MolmoSpaces, demonstrates capabilities including:
- Object picking and placement
- Opening drawers
- Operating doors
All tasks execute without training data from real-world demonstrations.
The Technical Approach
According to Ranjay Krishna, director of the PRIOR team at AI2, the key insight is straightforward: the simulation-to-reality gap shrinks dramatically when researchers increase the variety of simulated environments, objects, and camera conditions. Rather than improving physics simulation fidelity, the models benefit from diversity in training conditions.
This aligns with recent findings in robotics research showing that breadth of training distribution often matters more than pixel-perfect simulation accuracy. By exposing models to hundreds of thousands of variations in scene configuration, object types, and viewpoints, the models learn generalizable behaviors that transfer directly to physical systems.
Open-Source Release
Both models and supporting tools are available publicly. Technical details are available in the accompanying research paper. The open-source approach allows other research groups and robotics companies to build on the foundation rather than starting from zero with their own simulation data collection.
What This Means
This work addresses one of robotics' most significant friction points: the cost and time required to train deployable systems. If the zero-shot transfer results hold up in broader testing, the implications are substantial. Companies and research labs could dramatically reduce development timelines—from months of manual demonstration collection to weeks of model training. This could accelerate deployment of manipulation tasks in warehousing, manufacturing, and service robotics.
The emphasis on simulation diversity over physics accuracy also reframes how the robotics community should approach simulation tools. Rather than competing on fidelity, platforms that generate high-variance synthetic training data may prove more valuable. This could shift investment and resource allocation within the robotics software ecosystem.
Related Articles
Google DeepMind Launches Gemini Robotics 2, a Single VLA Model for Arms to Humanoids
Google DeepMind has introduced Gemini Robotics 2, a vision-language-action model it calls its most advanced yet, designed to control everything from tabletop robot arms to full-body humanoids. The company also released Gemini Robotics ER 2, an embodied reasoning model that replaces ER 1.6.
OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold
OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.
Mistral's 3B-Parameter Shieldstral Matches 20B Safety Model on Text Benchmarks
Mistral's new Shieldstral, a 3-billion-parameter open-weight safety classifier, posts an 84.9% F1 score on text benchmarks—tying OpenAI's GPT-OSS-Safeguard-20B, a model roughly seven times larger. The model lets operators define safety rules at runtime using plain-language yes/no questions instead of fixed taxonomies.
Mistral AI Releases Shieldstral-1.0-3B, a 3B-Parameter Policy-Adaptive Safety Classifier
Mistral AI has released Shieldstral-1.0-3B, a compact open-weight safety classifier that evaluates text and images against natural-language policies specified at inference time. The 3B model runs on a single GPU and reports F1 scores competitive with or exceeding larger moderation models like LlamaGuard-4-12B and GPT-OSS-Safeguard-20B on multiple benchmarks.
Comments
Loading...