NVIDIA Releases Cosmos3-Nano: 16B-Parameter Omnimodal World Model for Physical AI with 256K Token Context
NVIDIA has released Cosmos3-Nano, a 16-billion parameter omnimodal world model capable of generating video, audio, images, and robot action commands from combinations of text, image, video, and action trajectory inputs. The model supports a 256K token context window and is designed for Physical AI applications including robotics, autonomous vehicles, and smart manufacturing environments.
NVIDIA Releases Cosmos3-Nano: 16B-Parameter Omnimodal World Model for Physical AI
NVIDIA has released Cosmos3-Nano, a 16-billion parameter omnimodal world model that generates video, audio, images, and robot action commands from multimodal inputs. The model is part of the Cosmos3 collection and supports a 256K token context window for reasoning tasks.
Technical Specifications
Cosmos3-Nano uses a Mixture-of-Transformers (MoT) architecture combining an autoregressive transformer for discrete token generation with a diffusion transformer for continuous multimodal generation. The model accepts inputs in five modalities: text, images, video (with or without audio), and action trajectories.
Input specifications:
- Text: up to 256K tokens
- Images: 256p, 480p, and 720p at aspect ratios 16:9, 4:3, 1:1, 3:4, 9:16
- Video: same resolutions/aspect ratios, maximum 5 frames for input
- Audio: 48 kHz stereo, maximum 0.5 seconds
- Action trajectories: 16-400 frames
Output generation supports video from 5 to 400 frames (189 frames default), with audio encoded in AAC format at 48 kHz stereo. The model generates outputs at resolutions matching input specifications.
Model Variants
NVIDIA released multiple Cosmos3 variants simultaneously:
- Cosmos3-Nano: 16B parameters for general omnimodal tasks
- Cosmos3-Super: 64B parameters for enhanced performance
- Cosmos3-Nano-Policy-DROID: 16B parameters specialized for robot manipulation
- Cosmos3-Super-Image2Video: 64B parameters for image-to-video generation
- Cosmos3-Super-Text2Image: 64B parameters for text-to-image generation
Hardware and Deployment
The models require NVIDIA GPU-accelerated systems running on Ampere, Hopper, or Blackwell architectures. Only BF16 precision is officially supported. Runtime integration is available through PyTorch, vLLM-Omni, and Hugging Face Diffusers.
Cosmos3-Nano supports robot action generation for 10 different embodiments, including Franka Panda arms, WidowX 250, and various industrial robots. Action outputs are embodiment-specific, ranging from 9D to 57D depending on the robot platform.
Licensing and Availability
The model is released under the OpenMDW1.1 license for both commercial and non-commercial use. NVIDIA published the model collection on Hugging Face and GitHub on May 31, 2025, with global deployment availability.
Pricing information has not been disclosed.
What This Means
Cosmos3-Nano represents NVIDIA's entry into omnimodal foundation models specifically designed for embodied AI. The 256K context window for reasoning tasks and native support for action trajectory generation distinguishes it from general-purpose multimodal models. The architectural choice to combine autoregressive and diffusion transformers allows different generation mechanisms for discrete (text) versus continuous (video, audio, actions) modalities. With 10 supported robot embodiments at launch, NVIDIA is positioning Cosmos3 as infrastructure for physical AI development rather than a general-purpose model, directly targeting robotics researchers and autonomous system developers who need unified world modeling capabilities.
Related Articles
NVIDIA Releases Cosmos 3 Edge: 4B-Parameter World Model for Real-Time Robot Control at 15 Hz
NVIDIA has released Cosmos 3 Edge, a 4-billion-parameter open world model designed for edge AI systems. The model delivers real-time robot control at 15 Hz on NVIDIA Jetson devices, generating 32 actions per inference at 640×360 resolution.
Alibaba previews Qwen3.8 with 2.4 trillion parameters, claims second place without benchmark data
Alibaba unveiled Qwen3.8 at the World Artificial Intelligence Conference in Shanghai, claiming the 2.4 trillion parameter model ranks second only to Anthropic's Fable 5. The company provided no benchmark scores, model card, or independent verification to support the claim.
NVIDIA Releases Nemotron-3-Embed-1B-BF16: 1.14B Parameter Multilingual Embedding Model with 2048-Dimensional Vectors
NVIDIA has released Nemotron-3-Embed-1B-BF16, a 1.14 billion parameter text embedding model supporting 34 languages with a 32,768 token context window. The model generates 2048-dimensional embeddings and was derived from Ministral-3-3B-Instruct-2512 through two rounds of structured pruning and distillation, first to 2B then to 1.14B parameters.
Black Forest Labs releases FLUX.2: 32B open-weight image model with 4MP editing and 10-image multi-reference support
Black Forest Labs has released FLUX.2, a family of image generation models including a 32B parameter open-weight variant. The models support editing at up to 4 megapixel resolution and can reference up to 10 images simultaneously for character and style consistency.
Comments
Loading...