Back to Articles

We are excited to introduce Nemotron 3 Nano 4B, the newest and most compact member of the Nemotron 3 family. Leveraging hybrid Mamba-Transformer architecture, this model is designed for efficiency and accuracy in a targeted set of capabilities, setting a new standard for lightweight small language models. The model is available across any NVIDIA GPU-enabled platforms and combines state-of-the-art instruction following and exceptional tool use with minimal VRAM footprint.

With just 4 billion parameters, Nemotron 3 Nano 4B is compact enough to run at the edge on NVIDIA Jetson platforms (Jetson Thor/Jetson Orin Nano) as well as NVIDIA DGX Spark and NVIDIA RTX GPUs. This enables faster response times, enhanced data privacy, and flexible deployment while keeping inference costs low.

Nemotron 3 Nano 4B is our first model specifically optimized for on-device deployment and purpose-built to power local conversational agents and personas across GeForce RTX, Jetson and Spark customer use cases. This model achieves state-of-the-art accuracy and efficiency in several dimensions key to production use on the edge:

Instruction following (IFBench, IFEval): state-of-the-art in its size class