Nvidia released Nemotron 3.5 Lightning on Tuesday. It is a 30 billion parameter mixture-of-experts model with three billion active at any moment. The architecture is a hybrid of Mamba-2, MoE and attention layers, with a one million token context window. The weights are on Hugging Face and ModelScope.
The licence is the part that matters. It ships under OpenMDW-1.1, which permits commercial use, and Nvidia published the training data and the recipes alongside the weights. It is free for companies to download, use and modify without asking permission or paying Nvidia, CNBC notes.
The speed claim has two numbers
Nvidia leads on output speed of up to four times that of similar-sized models. Read further and the agentic figure is more modest. On PinchBench it hit 86% accuracy while finishing 10,000 tasks 30% faster than Alibaba’s Qwen 3.6 35B at similar accuracy.
Both numbers are in the same release. Four times faster describes token generation in the lab. Thirty percent is what happens when the model actually does a job. That gap is worth carrying whenever a vendor quotes throughput.










