World Labs Launches Atlas, an Omni-Model That Generates and Reconstructs 3D Worlds From a Handful of Photos
World Labs, co-founded by Fei-Fei Li, has released Atlas, an omni-model that generates camera-controlled video, reconstructs 3D scenes from as few as two images, and simulates environments for robot training. The company claims Atlas outperforms specialized 3D and video models on internal benchmarks, though it is currently limited to an early-access program.
World Labs, the startup co-founded by AI researcher Fei-Fei Li, has announced Atlas, a single model that generates, reconstructs, and simulates 3D scenes from as few as one to several dozen photos. The company claims Atlas beats specialized models at their own tasks — camera-controlled video generation, few-view 3D reconstruction, and robotics simulation — potentially making narrower tools unnecessary.
Atlas is World Labs' first model built around what the company calls "spatial intelligence." Unlike language or video-diffusion models that process data as flat one- or two-dimensional sequences, Atlas is trained from scratch on text, images, video, and 3D data, with every input anchored to a specific 3D position — a design World Labs calls "spatial context." Fei-Fei Li outlined this argument in a November 2025 essay, contending that current multimodal architectures make even simple spatial reasoning needlessly difficult.
Three core capabilities
For camera-controlled generation, Atlas accepts one or more images plus a camera path specified as geometric input — not a text description — and produces up to one minute of video at 1440p from any chosen viewpoint. World Labs says human evaluators preferred Atlas's output in head-to-head comparisons 75% of the time against MiniMax H3, 81% against Gemini Omni Flash, 86% against Happy Horse 1.1, and 94% against Seedance 2.5.
For spatial reconstruction, Atlas rebuilds real scenes from as few as two or three images and scales up to over a hundred inputs without specialized capture hardware. In one demo, it assembled Stanford's Main Quad from two to 25 ground-level photos and generated aerial views. World Labs reports a median reconstruction error of 25.3, ahead of Pi3X and VGGT-Ω 1B, under the OpenWorldLib evaluation framework, where competing systems like VGGT showed geometric inconsistencies once the camera moved significantly.
As a simulator, Atlas models space and time jointly, enabling "bullet time" effects from footage shot on ordinary smartphones and action cameras. It also serves as a real-to-sim engine for robotics, reconstructing a room and generating the sensor data — image and depth — a simulated robot would encounter, letting developers vary objects, lighting, and positions to generate training data without physical re-capture. This capability originated at SceniX, a startup World Labs acquired in July 2026, and World Labs previously showed a related real-to-sim-to-real engine in August 2026 running control models on five robot platforms for an hour each without human intervention, according to the company.
Atlas can also output native 3D formats — point clouds and 3D Gaussian splats — matching the representation used in World Labs' existing Marble product. Text-to-image generation is a secondary feature, supporting complex prompts, text rendering, and 360-degree panoramas.
Architecture and availability
World Labs describes Atlas as combining language-model-style piece-by-piece generation, which allows speedups like KV caching, with diffusion-based denoising for image quality. The company says performance improves with additional training compute and expects that trend to continue at scale. No parameter count, context window, or pricing has been disclosed.
Atlas is currently available only through an early-access program for select partners and will power future versions of Marble. World Labs, founded in 2024, raised a $1 billion round in February 2026 from Autodesk, Andreessen Horowitz, Nvidia, and AMD, following earlier reports of a $5 billion valuation.
What this means
Atlas is a bet that spatial reasoning deserves its own foundation model rather than being bolted onto video or language systems. The reconstruction and camera-control claims, if verified independently, would be meaningful for robotics simulation and 3D content pipelines that currently rely on stitching together multiple specialized tools. But the benchmarks come from World Labs itself, against a mix of named and unnamed competitors, and the model remains restricted to early-access partners — so broader validation, pricing, and general availability are still pending.
Related Articles
Anthropic Releases Claude Fable 5.1, Claims 52.6% on New Terminal-Bench-Science Benchmark
Anthropic released Claude Fable (and Mythos) 5.1, claiming a 52.6% score on the new Terminal-Bench-Science 0.1 benchmark — up sharply from 24.7% for Fable 5. Independent testing shows the model's five reasoning levels produce dramatically different output token counts and costs for identical prompts, ranging from $0.10 to $3.30 per request.
Meta Launches Muse Voice Transcribe, Real-Time Speech Model Handling 20+ Speakers Across 70+ Languages
Meta Superintelligence Lab released Muse Voice Transcribe, its first real-time audio perception model, claiming state-of-the-art streaming speech-to-text with native speaker diarization. The model handles 20+ speakers and code-switching across languages, priced at $3 per 1,000 audio minutes.
OpenAI's Astra Model Aces Cybersecurity Benchmark, Found Two Zero-Day Exploits Unassisted
OpenAI has disclosed new details on Astra, a forthcoming model the company says is the first to cross its 'critical cybersecurity threshold.' According to OpenAI, Astra scored a perfect result on ExploitBench and discovered two zero-day vulnerabilities in internal testing without human guidance.
OpenAI Says Upcoming Astra Model Is First to Cross 'Critical' Cybersecurity Risk Threshold
OpenAI says its upcoming Astra model is the first to cross its 'Critical' cybersecurity capability threshold, meaning it can discover and exploit unknown vulnerabilities without step-by-step human guidance. The company plans to release Astra soon but will restrict its advanced cyber capabilities to a vetted coalition of organizations.
Comments
Loading...