Its Vera Rubin NVL72 system can deliver up to 30 times the agentic-AI throughput per megawatt of its current GB300 NVL72 platform. The claim focuses on long-running agents, whose growing context, tool calls and sub-agents create a different load from a single chat response.

The company measured performance with SemiAnalysis AgentX, a workload built from recorded coding-agent sessions. Nvidia says Vera Rubin reached the largest advantage on DeepSeek V4 Pro and could lower cost per million tokens by as much as 35 times. Its DSX power-management software is also designed to fit more GPUs within a fixed power budget.

These are early vendor measurements, and Nvidia notes that the results are pending SemiAnalysis review. They also do not yet include Vera CPU performance for tool calls. The figures therefore should not be read as a universal gain for every model. The useful signal is the metric itself: as electricity and data-center capacity constrain deployments, operators are increasingly comparing complete agent workflows per megawatt rather than isolated chip speed.