Nvidia releases Nemotron 3 Super: 120B MoE model with 1M token context
Nvidia has released Nemotron 3 Super, a 120-billion parameter hybrid Mamba-Transformer Mixture-of-Experts model that activates only 12 billion parameters during inference. The open-weight model features a 1-million token context window, multi-token prediction capabilities, and pricing at $0.10 per million input tokens and $0.50 per million output tokens.
Nemotron 3 Super — Quick Specs
Nvidia Releases Nemotron 3 Super: 120B MoE Model with 1M Context Window
Nvidia has released Nemotron 3 Super, a 120-billion parameter open-weight model designed for multi-agent applications and long-context reasoning tasks. The model activates only 12 billion parameters during inference through a hybrid Mixture-of-Experts (MoE) architecture, balancing parameter scale with computational efficiency.
Key Specifications
Model Architecture: The model combines a Mamba-Transformer hybrid backbone with Mixture-of-Experts routing and multi-token prediction (MTP). Nvidia claims this design enables "over 50% higher token generation" compared to leading open-source models, though independent benchmarks are not yet available.
Context Window: 1 million tokens—one of the largest context windows available in open-weight models, enabling document analysis, cross-reference reasoning, and extended conversation memory.
Pricing: $0.10 per million input tokens and $0.50 per million output tokens via OpenRouter. This places it in the mid-tier pricing for large models, significantly cheaper than frontier closed-source alternatives.
Latent MoE Design: The model routes queries to 4 experts but applies computational cost equivalent to activating only one. Nvidia positions this as enabling "intelligence and generalization" improvements without proportional compute overhead.
Training and Performance
Nemotron 3 Super underwent multi-environment reinforcement learning training across 10+ simulation environments. According to Nvidia, the model achieves leading accuracy on AIME 2025, TerminalBench, and SWE-Bench Verified benchmarks. Specific benchmark scores have not been disclosed.
The model was released on March 11, 2026.
Licensing and Deployment
Nvidia released the model with full weights, training datasets, and recipes under the NVIDIA Open License. This enables customization and local deployment without cloud dependencies.
Market Position
Nemotron 3 Super enters a competitive space for open-weight reasoning models. The combination of 1M context window, MoE efficiency, and sub-$1 output token pricing targets developers building agentic systems and applications requiring extended context reasoning. The model's latent MoE approach represents an alternative to dense scaling or standard sparse MoE designs, though real-world efficiency gains require vendor-specific inference optimization.
Availability appears limited to OpenRouter as a primary provider, with additional routing partners handling fallback load.
What This Means
Nvidia's release signals continued commitment to open-weight models as strategic infrastructure for AI ecosystem players. The 1M context window and sub-12B activation pattern address two key pain points: expensive long-context reasoning and compute constraints in production deployments. However, performance claims lack independent verification, and real-world token generation speedup depends heavily on inference engine optimization—not guaranteed across all providers or hardware.
Related Articles
LG AI Research Releases K-EXAONE 2.0, a 750B-Parameter Open-Weight MoE Model with 262K Context
LG AI Research has released K-EXAONE 2.0, a 750-billion-parameter mixture-of-experts language model with 37B active parameters, a 262,144-token context window, and support for 10 languages. The model is open-weighted under Apache 2.0 and claims competitive results against Qwen3.5, GLM-5.1, and DeepSeek-V4 Pro on reasoning, coding, and long-context benchmarks.
NVIDIA Releases Nemotron VoiceChat 11B, an Open Full-Duplex Speech Model with Live Tool Calling
NVIDIA has released NemotronLabs VoiceChat 11B, an 11-billion-parameter end-to-end full-duplex speech model that unifies streaming speech understanding and generation in one architecture. The model claims to be the first open full-duplex system to support live tool calling during natural conversation, with ~450ms turn-taking latency.
Alibaba Unveils Qwen3.8-Max, a 2.4T-Parameter Open-Weight Model for Coding and Agentic Work
Alibaba announced Qwen3.8-Max, a 2.4T-parameter flagship model targeting coding and long-horizon agentic work, with open weights promised for next week alongside Qwen3.8-27B. The model posted strong third-party benchmark results, ranking #4 in Frontend Code Arena and matching Claude Opus 4.7 on the Vals Index at roughly 2.3x lower cost.
OpenAI Halts Parts of Astra Model Development After It Hit 'Critical' Cybersecurity Threshold
OpenAI disclosed that its in-development Astra model showed cyberattack capabilities strong enough that it cannot rule out a 'Critical' risk classification. The company has paused related internal activity and added security controls under its Preparedness Framework.
Comments
Loading...