Nvidia is handing out AI models like free samples at Costco. The logic is the same, too: once you try the product, you’ll come back to buy the hardware that runs it best.

The company’s latest release, Nemotron 3.5 Lightning, is a 30-billion-parameter model that launched on August 11, 2026. It uses a Mixture of Experts (MoE) architecture with roughly 3 billion active parameters, meaning it can run on a single GPU while supporting context windows up to 1 million tokens. That’s a serious amount of capability for a model you can download from Hugging Face without paying a dime.

The razor-and-blades playbook, supercharged

Nvidia now provides free hosted inference for over 100 AI models through its build.nvidia.com APIs. Developers can access these models without swiping a credit card, test them against their own workloads, and build applications on top of them.

Earlier in 2026, Nvidia dropped the Nemotron 3 Ultra, a 550-billion-parameter behemoth, alongside Dynamo 1.0, software that reportedly enhances GPU performance by up to 7x.