Nemotron 3.5 Lightning

Released
2026-08-11
Last refreshed
2026-10-05
Status
Researched 1d ago
Open weightsCommercial use: permittedRAGAgentsLong contextJSON / Tool use

Nemotron 3.5 Lightning is available now for rag, agents, and long context with open-weight and 1.05m context; evaluate it while provider pricing coverage matures.

NVIDIA AI releases · 20 in the last 12 months · this family litChangelog →
Specifications
Released
2026-08-11
Context
1.05m
Parameters
30B (3B activated)
Architecture
MoE + SSM Hybrid
Specialization
general
Openness
Open weights
License
OpenMDW 1.1Commercial use: permitted
Weights
Available
Code
Unknown
Training
Multi-stage
Created by

Accelerated AI for enterprise solutions

Santa Clara, California, United States
Founded 2015
Website
Pricing

No tracked provider token pricing is available yet.

About

NVIDIA Nemotron 3.5 Lightning (HF nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B) is an open-weight hybrid Mamba-Transformer mixture-of-experts text model with 30B total and 3B active parameters (NemotronHForCausalLM), released 11 August 2026. NVIDIA lists context up to 1M tokens (256K used for single-H100 deployment). Weights ship in BF16, NVFP4 and speculative-decoding (DSpark/DFlash) variants plus a base checkpoint under OpenMDW 1.1. It is offered on build.nvidia.com and OpenRouter (free tier).

Provider price ladder

No tracked provider token pricing is available for this model yet.

Capabilities

ReasoningJSON / Tool use

Benchmark peer barsfor RAG

No task-mapped benchmark peers are available for this model yet.

Migration checks

No linked migration route is available for this model yet.

API versions

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4NVIDIA-Nemotron-3.5-Lightning-30B-A3Bnvidia/nemotron-3.5-lightning