NVIDIA

GPU maker and AI infrastructure provider

https://nvidia.com

News

model releaseNVIDIA

NVIDIA Releases Nemotron-3-Embed-1B-BF16: 1.14B Parameter Multilingual Embedding Model with 2048-Dimensional Vectors

NVIDIA has released Nemotron-3-Embed-1B-BF16, a 1.14 billion parameter text embedding model supporting 34 languages with a 32,768 token context window. The model generates 2048-dimensional embeddings and was derived from Ministral-3-3B-Instruct-2512 through two rounds of structured pruning and distillation, first to 2B then to 1.14B parameters.

2 min read
product updateNVIDIA

NVIDIA NeMo Automodel integrates with Hugging Face Diffusers for distributed video and image model fine-tuning

NVIDIA and Hugging Face have integrated NeMo Automodel with the Diffusers library, enabling distributed fine-tuning of video and image diffusion models without checkpoint conversion. The integration supports models including FLUX.1-dev (12B), Wan 2.1 (1.3B/14B), and HunyuanVideo (13B) with full fine-tuning and LoRA options.

2 min read
model releaseNVIDIA

NVIDIA releases Nemotron-Labs-3-Puzzle-75B, compressed from 120B to 75B parameters with 2× throughput

NVIDIA has released Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super that reduces the model from 120.7B total/12.8B active parameters to 75.3B total/9.3B active parameters. According to NVIDIA, the model achieves approximately 2× higher server throughput on a single 8×B200 node and increases sustainable 1M-token single-H100 concurrency from 1 request to 8 requests while maintaining strong accuracy across benchmarks.

2 min read
researchNVIDIA

NVIDIA Releases 10 Trillion Tokens of Open Agentic Training Data, Launches Interactive Prompt Atlas

NVIDIA has released over 10 trillion pre-training tokens and millions of post-training samples as part of its Nemotron open data initiative for building AI agents. The release includes the Nemotron Post-Training v3 Prompt Atlas, an interactive visualization tool, and Nemotron-Personas dataset representing 2.4 billion people across 10 countries.

3 min read
model releaseNVIDIA

NVIDIA releases Nemotron-Labs-TwoTower-30B: block-wise diffusion model claims 2.42× faster generation at 98.7% baseline

NVIDIA released Nemotron-Labs-TwoTower-30B-A3B-Base-BF16, a block-wise diffusion language model that generates text by denoising blocks of tokens in parallel rather than sequentially. According to NVIDIA, the model achieves 2.42× the wall-clock generation throughput of its autoregressive baseline while retaining 98.7% of aggregate benchmark quality.

2 min read
product updateNVIDIA

AWS brings NVIDIA Nemotron and OpenAI GPT OSS models to GovCloud for secure government AI workloads

Amazon Bedrock now supports NVIDIA Nemotron and OpenAI GPT OSS models in AWS GovCloud (US) Regions. The launch includes OpenAI's GPT OSS models (120B and 20B parameters, 128K context) and NVIDIA Nemotron 3 family (9B to 120B parameters, 1M context), providing government agencies FedRAMP High and DoD SRG Level 5-compliant AI inference on U.S. soil.

2 min read
model releaseNVIDIA

Nvidia releases Nemotron 3 Ultra: 550B-parameter MoE model with 1M context window for agentic workflows

Nvidia has released Nemotron 3 Ultra, a 550-billion parameter mixture-of-experts model with 55 billion active parameters and support for up to 1 million token context windows. The model uses a hybrid Transformer-Mamba architecture and is designed specifically for long-running agentic workflows including agent orchestration, coding agents, and complex enterprise tasks.

2 min read
model releaseNVIDIA

NVIDIA Releases Nemotron-3-Ultra: 550B Parameter Model with 1M Token Context and Configurable Reasoning

NVIDIA released Nemotron-3-Ultra-550B-A55B-NVFP4, a 550B parameter model with 55B active parameters, featuring a 1M token context window and configurable reasoning mode. The model uses a hybrid LatentMoE architecture combining Mamba-2, Mixture-of-Experts, and Attention layers with Multi-Token Prediction, trained with NVIDIA's NVFP4 quantization-aware approach.

2 min read
model releaseNVIDIA

NVIDIA releases Nemotron-3-Ultra: 550B parameter model with 1M token context and configurable reasoning

NVIDIA released Nemotron-3-Ultra-550B, a frontier-scale model with 550B total parameters (55B active) and up to 1M token context window. The model uses a hybrid LatentMoE architecture combining Mamba-2, MoE, and attention layers with Multi-Token Prediction, trained with NVFP4 quantization-aware methods from December 2025 to April 2026.

2 min read
model releaseNVIDIA

NVIDIA Releases Nemotron 3.5 Content Safety: 4B-Parameter Multimodal Model with Custom Policy Enforcement and 140-Langua

NVIDIA has released Nemotron 3.5 Content Safety, a 4B-parameter model built on Google Gemma 3 4B IT that provides multimodal safety classification across approximately 140 languages. The model includes a 128K context window, custom enterprise policy enforcement, auditable reasoning traces, and is releasing its training dataset.

3 min read

Models

Nemotron-3-Embed-1B-BF16

NVIDIA

active
Context33K

Jul 20, 2026

Cosmos 3 Edge

NVIDIA

active

Jul 16, 2026

Nemotron 3 Embed 8B BF16

NVIDIA

active
Context32K

Jul 16, 2026

Audex-30B-A3B

NVIDIA

active
Context1000K

Jul 9, 2026

NVIDIA Nemotron-Labs-3-Puzzle-75B-A9B-NVFP4

NVIDIA

active
Context1000K

Jul 9, 2026

Nemotron-Labs-TwoTower-30B-A3B-Base-BF16

NVIDIA

active

Jul 4, 2026

DiffusionGemma 26B A4B IT NVFP4

NVIDIA

active
Context262K

Jun 17, 2026

Nemotron 3 Ultra

NVIDIA

active
Context1000K
Input/1M$0.5

Jun 5, 2026

Nemotron 3.5 ASR

NVIDIA

active

Jun 4, 2026

Nemotron 3.5 Content Safety

NVIDIA

active
Context128K
0

Jun 4, 2026

Nemotron-3-Ultra-550B-A55B

NVIDIA

active
Context1000K

Jun 4, 2026

Cosmos3-Nano

NVIDIA

active
Context256K

Jun 2, 2026

Cosmos 3 Super

NVIDIA

active
Context256K

Jun 1, 2026

Cosmos 3 Super Image2Video

NVIDIA

active
Context262K

May 31, 2026

Cosmos3-Super-Text2Image

NVIDIA

active
Context4K

May 31, 2026

LocateAnything-3B

NVIDIA

active
Context24K

May 26, 2026

Nemotron-Labs Diffusion 8B

NVIDIA

active

May 23, 2026

Nemotron 3 Nano Omni 30B-A3B-Reasoning

NVIDIA

active
Context256K

Apr 28, 2026

Nemotron-3-Nano-Omni-30B-A3B

NVIDIA

active
Context256K

Apr 28, 2026

NVIDIA Isaac GR00T N1.7

NVIDIA

active

Apr 17, 2026

Gemma 4 31B IT NVFP4

NVIDIA

active
Context262K

Apr 2, 2026

gpt-oss-puzzle-88B

NVIDIA

active
Context128K

Mar 26, 2026

NVIDIA Nemotron-3-Nano-4B-GGUF

NVIDIA

active
Context262K

Mar 16, 2026

Nemotron 3 Super

NVIDIA

active
Context1000K
Input/1M$0.1

Mar 11, 2026

NVIDIA Nemotron-3-Super-120B-A12B

NVIDIA

active
Context1000K
Input/1M$0.2

Mar 10, 2026

Nemotron 3 Content Safety 4B

NVIDIA

active
Context128K

Mar 20, 2025

Llama 3.1 Nemotron 70B Instruct

NVIDIA

active
Context128K
Input/1M$0.2

Oct 15, 2024