NVIDIA Releases Cosmos3-Nano: 16B-Parameter Omnimodal World Model for Physical AI with 256K Token Context
NVIDIA has released Cosmos3-Nano, a 16-billion parameter omnimodal world model capable of generating video, audio, images, and robot action commands from combinations of text, image, video, and action trajectory inputs. The model supports a 256K token context window and is designed for Physical AI applications including robotics, autonomous vehicles, and smart manufacturing environments.
NVIDIA Releases Cosmos3-Nano: 16B-Parameter Omnimodal World Model for Physical AI
NVIDIA has released Cosmos3-Nano, a 16-billion parameter omnimodal world model that generates video, audio, images, and robot action commands from multimodal inputs. The model is part of the Cosmos3 collection and supports a 256K token context window for reasoning tasks.
Technical Specifications
Cosmos3-Nano uses a Mixture-of-Transformers (MoT) architecture combining an autoregressive transformer for discrete token generation with a diffusion transformer for continuous multimodal generation. The model accepts inputs in five modalities: text, images, video (with or without audio), and action trajectories.
Input specifications:
- Text: up to 256K tokens
- Images: 256p, 480p, and 720p at aspect ratios 16:9, 4:3, 1:1, 3:4, 9:16
- Video: same resolutions/aspect ratios, maximum 5 frames for input
- Audio: 48 kHz stereo, maximum 0.5 seconds
- Action trajectories: 16-400 frames
Output generation supports video from 5 to 400 frames (189 frames default), with audio encoded in AAC format at 48 kHz stereo. The model generates outputs at resolutions matching input specifications.
Model Variants
NVIDIA released multiple Cosmos3 variants simultaneously:
- Cosmos3-Nano: 16B parameters for general omnimodal tasks
- Cosmos3-Super: 64B parameters for enhanced performance
- Cosmos3-Nano-Policy-DROID: 16B parameters specialized for robot manipulation
- Cosmos3-Super-Image2Video: 64B parameters for image-to-video generation
- Cosmos3-Super-Text2Image: 64B parameters for text-to-image generation
Hardware and Deployment
The models require NVIDIA GPU-accelerated systems running on Ampere, Hopper, or Blackwell architectures. Only BF16 precision is officially supported. Runtime integration is available through PyTorch, vLLM-Omni, and Hugging Face Diffusers.
Cosmos3-Nano supports robot action generation for 10 different embodiments, including Franka Panda arms, WidowX 250, and various industrial robots. Action outputs are embodiment-specific, ranging from 9D to 57D depending on the robot platform.
Licensing and Availability
The model is released under the OpenMDW1.1 license for both commercial and non-commercial use. NVIDIA published the model collection on Hugging Face and GitHub on May 31, 2025, with global deployment availability.
Pricing information has not been disclosed.
What This Means
Cosmos3-Nano represents NVIDIA's entry into omnimodal foundation models specifically designed for embodied AI. The 256K context window for reasoning tasks and native support for action trajectory generation distinguishes it from general-purpose multimodal models. The architectural choice to combine autoregressive and diffusion transformers allows different generation mechanisms for discrete (text) versus continuous (video, audio, actions) modalities. With 10 supported robot embodiments at launch, NVIDIA is positioning Cosmos3 as infrastructure for physical AI development rather than a general-purpose model, directly targeting robotics researchers and autonomous system developers who need unified world modeling capabilities.
Related Articles
GLM-5.3-Flash Debuts as Zhipu AI's First Natively Multimodal Model, 320B Parameters with 18B Active
Zhipu AI has released GLM-5.3-Flash, the first natively multimodal model in its GLM-5 series, built on a 320B-parameter mixture-of-experts architecture with only 18B active parameters. The company claims it outperforms GLM-5.2 while approaching Claude Opus 4.8 on coding and agentic benchmarks at a fraction of the cost. Unsloth has published quantized GGUF versions for local inference.
Z.ai Launches GLM-5.3-Flash: 1M-Token Context, Image Support, Claimed 10x Cost Cut Over GLM-5.2
Z.ai has released GLM-5.3-Flash, a 320-billion-parameter Mixture-of-Experts model with 18 billion active parameters, a 1-million-token context window, and image input support. The model launched on LM Studio's Bionic platform hours after its official unveiling, with LM Studio claiming it is 9-10x cheaper to run than GLM-5.2.
Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context
Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.
Z.ai Launches GLM-5.3-Flash With 1M-Token Context and Hybrid Attention Architecture
Z.ai has released GLM-5.3-Flash, a native multimodal model built for coding and long-horizon agent tasks, featuring a 1M-token context window and a hybrid sparse-linear attention architecture. The model is available via OpenRouter at a discounted $0.075/$0.25 per 1M tokens through September 2026.
Comments
Loading...