Long-running AI agents spend most of their time on high-volume execution: tool calls, result validation, and subagent delegation. Using a frontier reasoning model for every execution step adds cost…

Nemotron 3.5 Lightning is designed to work alongside other models in AI agents, targeting specific jobs like code review, answering billing questions, and security alert…

Nvidia's Nemotron 3.5 Lightning pairs with its NeMo Switchyard router, which reassigns models mid-task and cuts task costs to a third in Nvidia's own tests.