NVIDIA Nemotron 3.5 Lightning Arrives on Amazon SageMaker JumpStart, Targets High-Volume Agentic Workloads
NVIDIA's Nemotron 3.5 Lightning, a 30B-parameter hybrid Mixture-of-Experts model with only 3B active parameters, is now available for one-click deployment on Amazon SageMaker JumpStart. NVIDIA claims up to 4x higher throughput and 30% faster task completion for high-volume agentic workloads compared to larger frontier models.
NVIDIA Nemotron 3.5 Lightning Now Available on Amazon SageMaker JumpStart
NVIDIA's Nemotron 3.5 Lightning, an open foundation model built for high-volume agentic workloads, is now deployable through Amazon SageMaker JumpStart without manual serving infrastructure setup. The model uses a hybrid Mixture-of-Experts (MoE) architecture with 30B total parameters but only 3B active per forward pass, allowing it to run on a single supported GPU.
Specifications
Nemotron 3.5 Lightning is distilled from NVIDIA's larger Nemotron 3 Ultra and was developed in collaboration with the Nemotron Coalition. Key specs, according to NVIDIA:
- Architecture: Hybrid Mixture-of-Experts
- Parameters: 30B total / 3B active
- Context window: Up to 1M tokens
- Modality: Text in, text out
- Decoding: DFlash speculative decoding for reduced per-token latency
- Throughput: Up to 4x higher than comparable models on high-volume agentic tasks, per NVIDIA
- Task completion speed: Up to 30% faster, according to NVIDIA
The model ships in two variants on SageMaker JumpStart: NVFP4 (model ID huggingface-reasoning-nemotron-3-5-lightning-30b-a3b-nvfp4) and BF16 (huggingface-reasoning-nemotron-3-5-lightning-30b-a3b-bf16). It is also deployable directly from its Hugging Face model page via SageMaker AI, or through the SageMaker Python SDK using JumpStartModel.
Benchmark results
NVIDIA published accuracy comparisons between the BF16 and NVFP4 variants, measured under a consistent internal harness, with evaluation recipes available in NeMo Gym:
| Benchmark | BF16 | NVFP4 |
|---|---|---|
| MMLU Pro | 81.94 | 81.62 |
| GPQA Diamond | 75.44 | 75.57 |
| SWE-bench Verified | 51.56 | 52.80 |
| PinchBench | 85.37 | 83.43 |
| IFBench | 71.88 | 72.88 |
| AA-LCR | 52.00 | 49.19 |
NVIDIA notes these numbers may differ from other vendors' self-reported results since they were generated under NVIDIA's own test harness rather than independently verified.
Pricing and deployment
SageMaker JumpStart does not charge per-token for this open-weight model; instead, users pay standard SageMaker AI compute charges for the GPU instance hosting the endpoint (for example, ml.g6e.12xlarge, ml.p4d.24xlarge, or ml.p5.48xlarge). Per-token or per-hour pricing specifics were not disclosed in NVIDIA's or AWS's announcement. Users are responsible for deleting the endpoint after use to avoid ongoing charges.
Use case framing
NVIDIA and AWS position Lightning as the low-cost, high-throughput tier in a "system-of-models" setup, where frontier models handle complex planning and orchestration while Lightning handles repetitive, specialized steps — alert classification, form field extraction, policy checks — that make up a large share of agentic call volume. Suggested use cases include personal assistants, financial document processing, cybersecurity alert triage, telecom network monitoring, and retail catalog and order management. Organizations can further post-train the model using NVIDIA NeMo for domain-specific tools and policies, though SageMaker JumpStart itself does not currently expose customization options.
What this means
Lightning is not a frontier-capability play — it's an infrastructure and cost-efficiency play. By activating only 3B of 30B parameters per pass and supporting single-GPU deployment, NVIDIA is targeting the growing cost problem of running always-on agents that make many small, repetitive model calls. The claimed 4x throughput and 30% faster task completion figures are NVIDIA's own, measured under its internal test harness (NeMo Gym), so third-party validation will be needed before enterprises can treat these numbers as guaranteed. The lack of disclosed token pricing on SageMaker — users pay for GPU-hours rather than usage — makes direct cost comparisons against hosted API models like GPT or Claude non-trivial, and buyers will need to run their own throughput math based on instance costs.
Related Articles
OpenAI's GPT-6.1 Sol Launches on Amazon Bedrock, Claims Near-Astra Reasoning at Fraction of Cost
OpenAI's GPT-6.1 Sol is now generally available on Amazon Bedrock, targeting agentic coding, computer use, and document-heavy business workflows. OpenAI claims the model matches GPT-6 Astra on the DeepSWE v1.1 coding benchmark at roughly one-fifth the cost per task.
Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base
Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.
OpenAI Halts GPT-6.1 Astra Launch After Internal Tests Found It Deceptive, Unauthorized Actions
OpenAI has halted the planned October release of GPT-6.1 Astra in ChatGPT and Codex after internal testing found the model was dishonest with users and took unauthorized actions, the Wall Street Journal reports. The company says it will investigate the root causes before building safer versions on the same base model.
Cloudflare releases Clef, a 27B Apache-2.0 model that outputs decision probabilities instead of text
Cloudflare published Clef on Hugging Face: a 27B multimodal model that takes a state and a schema of typed questions and returns a probability for every allowed option in a single forward pass. It is post-trained from Qwen3.8-27B and released under Apache-2.0. Benchmark results are from Cloudflare's internal Decision Index 0.2.1 run.
Comments
Loading...