
As Anticipation for Rubin Builds, NVIDIA Continues to Optimize Blackwell, Quadrupling GB200 Energy Efficiency in Three Months
Through 38 major software optimizations and 250,000 tests, NVIDIA increased the energy efficiency of running the DeepSeek model on the GB200 by fourfold in just three months. Meanwhile, the GB300 NVL72 set new training records, with performance growing 1.5 times over six months and scaling efficiency nearing 98.5%. As the new Rubin architecture accelerates its deployment, NVIDIA continues to unlock the potential of the Blackwell platform, demonstrating strong hardware-software synergy
NVIDIA is extending the lifecycle of the Blackwell platform through continuous software stack optimization, while simultaneously building momentum for the large-scale deployment of its next-generation Rubin architecture.
Against the backdrop of accelerated deployment of the new Rubin platform, NVIDIA disclosed that it quadrupled the compute throughput per megawatt (TPS/MW) of the GB200 NVL72 when running the DeepSeek R1 0528 model in just three months. This achievement stems from 38 major optimizations completed by NVIDIA over four months, supported by more than 250,000 simulated configuration tests and a cumulative 1.4 million GPU hours of optimization verification.
Meanwhile, the GB300 "Blackwell Ultra" platform continues to break records in AI training benchmarks. At a scale of 256 GPUs, the GB300 NVL72 achieved a historical high of 1,648 TFLOPs per GPU in the DeepSeek-V3 671B pre-training task, approximately triple the 606 TFLOPs of the GB200. This metric has cumulatively increased by 1.5 times over the past six months.
Significant Leap in GB200 Energy Efficiency, Optimizations Cover Over 90% of Models
NVIDIA stated that a series of optimizations for the GB200 NVL72 platform resulted in a fourfold increase in compute throughput per megawatt within three months when running DeepSeek R1 0528 (with a 1K input/1K output configuration).

Supporting this improvement were 38 major optimization iterations completed by NVIDIA over four months. These optimizations were screened through more than 250,000 simulated configurations and verified through actual testing consuming a cumulative 1.4 million GPU hours. NVIDIA specifically noted that over 90% of these optimization results are reusable across models, applicable to other models in its AI product portfolio, rather than being tailor-made for a single task.
This means existing data center customers can achieve significant energy efficiency gains through software upgrades without replacing hardware—a direct economic benefit for hyperscale cloud providers and enterprise customers who are highly sensitive to compute costs.
GB300 Breaks Training Records, Performance Increases 1.5 Times in Six Months
On the AI training front, the GB300 NVL72 also demonstrates strong performance momentum. At a scale of 256 GPUs, pre-training the DeepSeek-V3 671B model based on the Megatron Core framework, the GB300 NVL72 achieved a throughput of 1,648 TFLOPs per GPU, approximately triple the previous 606 TFLOPs of the GB200, setting a new global record for this task.

Notably, this figure does not represent a static hardware limit. The performance of the GB300 NVL72 under the Megatron Core framework has grown from 1,088 TFLOPs/GPU in November 2025 to 1,648 TFLOPs/GPU in June 2026, a cumulative increase of approximately 1.5 times over six months.

In terms of collaborative optimization with mainstream AI frameworks, NVIDIA's deep cooperation with the PyTorch and JAX communities has also brought significant gains. On TorchTitan (PyTorch's native training stack), the training performance of the GB300 NVL72 on DeepSeek-V3 671B improved sixfold compared to the unoptimized baseline configuration, jumping from 199 TFLOPs/GPU to 1,197 TFLOPs/GPU. The improvement under the JAX framework was even more pronounced; as of July 2026, throughput per GPU reached 4,082 Tokens/s, corresponding to 1,025 TFLOPs/GPU, an approximately tenfold increase from 418 Tokens/s in January 2026.

Scaling Efficiency Nears Theoretical Limit, 800 Gb/s Network Is Key
Scaling efficiency in large-scale training scenarios has always been one of the core metrics for measuring the practicality of AI infrastructure. NVIDIA disclosed that within the scaling range of 256 to 1,024 GPUs, the GB300 NVL72 maintained scaling efficiency close to the theoretical limit across three mainstream frameworks: 98.5% for Megatron Core, and 97% for both TorchTitan and JAX.
NVIDIA attributed this scaling performance to the 800 Gb/s Scale-Out network chips built into the NVL72 rack. High-speed interconnects directly determine communication overhead in multi-node training, thereby impacting overall scaling efficiency. This networking capability is regarded as the core infrastructure component supporting the aforementioned high-efficiency performance across frameworks.

Rubin Platform Accelerates Deployment, Blackwell Optimization Continues
As these optimization advances were announced, NVIDIA's next-generation Vera Rubin platform had already entered the global deployment phase. Reportedly, the Vera Rubin NVL72 achieves approximately a 10x increase in token throughput compared to Blackwell, with the GB200 NVL72 at around 80,000 Tokens/s, while the Vera Rubin NVL72 can reach 800,000 Tokens/s under the same 150MW power consumption.
However, with the Blackwell platform already deployed in countless data centers globally, NVIDIA's strategy mirrors that of the previous Hopper generation—continuously unlocking the potential of deployed hardware through software optimization while advancing the new platform. This approach not only extends the return on investment period for existing customers' hardware but also strengthens NVIDIA's platform stickiness within the AI infrastructure ecosystem. For the market, NVIDIA's dual-track advancement with Blackwell and Rubin is gradually building a software moat that competitors will find difficult to surpass in the short term.
