As Rubin is highly anticipated, Nvidia continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months.

As Rubin is highly anticipated, Nvidia continues to optimize Blackwell, with GB200's energy efficiency quadrupling in three months.

```

NVIDIA is extending the lifecycle of the Blackwell platform through continuous software stack optimization, while preparing for the large-scale rollout of the next-generation Rubin architecture.

Against the backdrop of accelerating deployment of the new generation Rubin platform, NVIDIA revealed that in just three months, it increased the GB200 NVL72’s power throughput per megawatt (TPS/MW) when running the DeepSeek R1 0528 model by four times. This achievement stems from 38 major optimizations completed within four months, backed by more than 250,000 simulated configuration tests and a cumulative 1.4 million GPU hours of validation.

Meanwhile, the GB300 "Blackwell Ultra" platform continues to break records in AI training benchmarks. At a scale of 256 cards, GB300 NVL72 achieved a historic record in the DeepSeek-V3 671B pre-training task with 1648 TFLOPs per GPU, about three times higher than GB200’s 606 TFLOPs, and this metric has increased 1.5 times in the past six months.

GB200 Power Efficiency Leaps—Optimizations Cover Over 90% of Models

NVIDIA stated that a series of optimizations for the GB200 NVL72 platform enabled a four-fold increase in power throughput per megawatt within three months when running DeepSeek R1 0528 (1K input/1K output configuration).

The boost is supported by 38 major optimization iterations completed within four months. These optimizations underwent more than 250,000 simulated configuration screenings and consumed a total of 1.4 million GPU hours for validation. NVIDIA highlighted that over 90% of the optimization results are reusable across models and suitable for other models in its AI product portfolio, rather than tailored for single tasks.

This means that existing data center customers can gain significant energy efficiency benefits through software updates without needing to replace hardware—a directly economic value proposition for hyperscale cloud providers and enterprise clients highly sensitive to compute costs.

GB300 Breaks Training Records—with Performance Up 1.5 Times in Six Months

On the AI training side, GB300 NVL72 also demonstrates strong performance momentum. At a scale of 256 cards, using the Megatron Core framework for pre-training the DeepSeek-V3 671B model, GB300 NVL72 achieved a throughput of 1648 TFLOPs per GPU, about three times higher than the previous GB200’s 606 TFLOPs, setting a global record for the task.

Notably, this number is not a static hardware limit. GB300 NVL72’s Megatron Core performance increased from 1088 TFLOPs/GPU in November 2025 to 1648 TFLOPs/GPU in June 2026, a cumulative improvement of about 1.5 times over six months.

In terms of collaborative optimization with mainstream AI frameworks, NVIDIA’s deep cooperation with the PyTorch and JAX communities also yielded significant gains. On TorchTitan (PyTorch native training stack), GB300 NVL72’s training performance for DeepSeek-V3 671B increased six times compared to the unoptimized baseline, jumping from 199 TFLOPs/GPU to 1197 TFLOPs/GPU. The improvement on the JAX framework is even more pronounced, reaching 4082 Tokens/s per GPU (corresponding to 1025 TFLOPs/GPU) as of July 2026, about ten times higher than the 418 Tokens/s in January 2026.

Expansion Efficiency Nears Theoretical Maximum—800 Gb/s Network Is Key

Expansion efficiency in large-scale training scenarios has always been a core indicator for assessing the practicality of AI infrastructure. NVIDIA disclosed that in the expansion range from 256 to 1024 cards, GB300 NVL72 maintained expansion efficiency close to the theoretical maximum across three major frameworks: Megatron Core reached 98.5%, and both TorchTitan and JAX reached 97%.

NVIDIA attributes this expansion performance to the NVL72 rack’s built-in 800 Gb/s scale-out network chip. High-speed interconnect directly determines communication overhead in multi-machine training, which in turn affects overall expansion efficiency; this network capability is viewed as the core infrastructure component supporting the high efficiency across frameworks.

Rubin Platform Rolls Out Faster, While Blackwell Optimization Continues

As these optimization advances are released, NVIDIA’s next-generation Vera Rubin platform has entered the global deployment phase. Reportedly, Vera Rubin NVL72 achieves about a ten-fold improvement in token throughput compared to Blackwell: GB200 NVL72 delivers about 80,000 Tokens/s, while Vera Rubin NVL72 reaches up to 800,000 Tokens/s under the same 150 MW power consumption.

However, the Blackwell platform has already been deployed in countless data centers worldwide, and NVIDIA’s strategy is similar to the previous Hopper generation—continuing to tap the potential of already-deployed hardware through ongoing software optimization as the new platform advances. This approach not only extends the ROI cycle of existing customer hardware investments but also strengthens NVIDIA’s platform stickiness in the AI infrastructure ecosystem. In the market, NVIDIA’s dual-track progress with Blackwell and Rubin is gradually building a software moat that competitors will find difficult to surpass in the short term.

Risk Warning and DisclaimerThe market has risks, and investment should be cautious. This article does not constitute personal investment advice, nor does it take into account the specific investment objectives, financial situations, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their particular situation. Responsibility for investing based on these lies solely with the user. ```