Nvidia’s Vera Rubin NVL72 system delivered up to 3.7 times the throughput of a GB300 NVL72 rack on Qwen3-VL in its first MLPerf Inference preview submission. On DeepSeek-R1, the new rack reached up to 2.5 times GB300 throughput, according to results released with MLPerf Inference v6.1.

The submissions used different software stacks for the two workloads: vLLM with Nvidia Dynamo for Qwen3-VL and TensorRT-LLM for DeepSeek-R1. Nvidia attributes the gains to updated Tensor Cores, its Transformer Engine, lower-precision NVFP4 processing and an NVLink fabric designed for higher packet rates and lower latency. The system separates prompt processing from token generation and distributes mixture-of-experts work across the rack.

A separate GB300 submission scaled DeepSeek-R1 from one 72-GPU rack to four racks, or 288 GPUs, with 99% scaling efficiency in the offline test. Software updates also produced gains of up to 1.6 times over the previous MLPerf round. The Vera Rubin numbers are preview results and vendor comparisons, so buyers should distinguish them from generally available production deployments and examine the precise benchmark configuration. Even so, the results indicate that Nvidia is pursuing inference gains through tightly coordinated hardware, networking and serving software rather than chips alone.