MLPerf Inference v6.1 Delivers First Preview Results for Vera Rubin NVL72
MLCommons published the MLPerf Inference v6.1 benchmark results. NVIDIA submitted the first preview results for the Vera Rubin NVL72 system for the DeepSeek-R1 and Qwen3-VL 235B models.

MLCommons published the results of the MLPerf Inference v6.1 benchmark on September 16, 2026. They include the first publicly released results for the Vera Rubin NVL72 system, which NVIDIA submitted in preview mode for the DeepSeek-R1 and Qwen3-VL 235B models.
MLPerf Inference measures the performance of AI systems during inference, meaning the operation of already-trained models. Measurements are conducted under defined conditions, intended to enable comparisons of the capacity and responsiveness of individual configurations in the same scenarios.
Vera Rubin NVL72 in interactive tests
In an interactive test with the Qwen3-VL 235B vision-language model, the system achieved 1,307 queries per second. NVIDIA compares it with the GB300 NVL72 configuration, which achieved 349 queries per second in the same test. This represents an approximately 3.7-fold difference in throughput.
For the DeepSeek-R1 model, which measures the number of generated tokens, Vera Rubin NVL72 achieved 652,750 tokens per second. GB300 NVL72 recorded 260,098 tokens per second, representing approximately 2.5 times the performance.
- Qwen3-VL 235B: 1,307 queries per second on the Vera Rubin NVL72 system versus 349 on the GB300 NVL72 configuration.
- DeepSeek-R1: 652,750 tokens per second on the Vera Rubin NVL72 system versus 260,098 on the GB300 NVL72 configuration.
The results are still only preview results
Both MLCommons and NVIDIA identify the submitted Vera Rubin NVL72 results as a preview submission. They are therefore not the platform’s final results in the MLPerf Inference v6.1 benchmark.
The figures also apply to the specific published configurations, models, and interactive scenarios. They therefore cannot be interpreted as a universal comparison of all AI accelerators or as a direct representation of performance in every real-world deployment.
The published results do not indicate the systems’ price, energy consumption, or customer availability date. NVIDIA also cites later optimizations, but MLCommons did not verify them as part of the v6.1 release, and they are not included in its official results.
What will matter next
For AI infrastructure operators, MLPerf results are one input when assessing system capacity for large language and vision-language models. For Vera Rubin NVL72, it will be important to see when final results, power-consumption and cost-per-token data, and comparisons across a broader set of models and competing configurations become available.
Sources
- MLCommons – Confirms the release of MLPerf Inference v6.1, the designation of NVIDIA Rubin and Vera Rubin NVL72 as preview results, and the context of the benchmark’s new results.
- NVIDIA Developer – Provides the specific measured values for Vera Rubin NVL72 and GB300 NVL72 in the DeepSeek-R1 and Qwen3-VL scenarios; it also explicitly identifies the Vera Rubin results as preview results.
- NVIDIA Blog – Documents NVIDIA’s announcement and its claims of up to 3.7 times the throughput for Qwen3-VL and 2.5 times for DeepSeek-R1.
Verified and updated: 09/17/2026 06:26



