NVIDIA Nemotron 3.5 Lightning 30B-A3B-NVFP4

NVIDIA🇺🇸 United States
active
Context window1000K tokens

Version History

3.5-Lightning-30B-A3B-NVFP4major

NVIDIA released a new NVFP4-quantized checkpoint of Nemotron 3.5 Lightning, a 30B-total/3B-active hybrid Mamba-2/MoE/Attention model with 1M-token context, optimized for single-GPU deployment on DGX Spark or H100 hardware. The release includes post-training quantization and speculative decoding support to preserve BF16-level accuracy at lower inference cost.

Coverage

model releaseNVIDIA

NVIDIA Releases Nemotron 3.5 Lightning 30B-A3B: 3B-Active MoE Model With 1M-Token Context, Quantized for Single-GPU Depl

NVIDIA has published NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4, a 30-billion-parameter Mixture-of-Experts model with only 3B active parameters, a hybrid Mamba-2/MoE/Attention architecture, and support for up to 1 million tokens of context. The NVFP4-quantized checkpoint is designed to run on a single DGX Spark (GB10) or H100 GPU.

2 min read