What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale. These factories are now tasked with powering agentic workflows that reason, plan, use tools, verify intermediate results, and execute complex multistep tasks across vast contexts.

Agentic workloads are not defined by a single prompt and response, but by sustained inference across many reasoning steps. They demand low per-step latency, high decode throughput, efficient long-context attention, large KV cache capacity, and the ability to scale models across tightly coupled GPU domains. The data center must be reimagined as a single unit of compute, a vision realized with the NVIDIA Vera Rubin platform.

At the core of the platform is the NVIDIA Rubin GPU, designed to deliver up to 10x more agentic throughput per unit of energy than NVIDIA Blackwell (Figure 1). Enhanced Tensor Cores with expanded precision flexibility, a new HBM4 memory subsystem, and the third-generation Transformer Engine—delivering up to 50 petaflops of sparse NVFP4 performance—work together to accelerate agentic workloads efficiently.

Figure 1. Pareto frontiers illustrating the 10x generational uplift in agentic inference performance of the Vera Rubin platform (internal 2T MoE workload)