Nvidia just proved that sometimes the best hardware upgrade is better software. The company revealed on July 1 that its optimized inference stack can slash token costs by up to 5x for DeepSeek V4 running on Blackwell systems, with throughput improvements reaching as high as 20x compared to baseline configurations on the exact same chips.
Those gains didn’t come from a new chip announcement or a fresh silicon architecture. They came from roughly one month of software engineering on open-source frameworks like vLLM and SGLang, plus a cocktail of techniques including disaggregated serving, NVLink expert parallelism, NVFP4 precision, and multi-token prediction.
What the numbers actually mean
Nvidia is increasingly pushing two metrics it wants the industry to obsess over: tokens per dollar and tokens per watt. Both measure how much useful AI output you squeeze from a given investment in hardware and electricity.
Enterprise inference provider Baseten, collaborating with Nvidia on the rollout, demonstrated a 50% increase in tokens per second using TensorRT-LLM on DeepSeek V4 Pro systems. That’s a more conservative number than the headline figures, but it reflects real-world production conditions rather than lab benchmarks.








