NVIDIA's Blackwell GB300 and GB200 GPUs continue to showcase leading capabilities across AI workloads through continuous optimizations.
NVIDIA Isn't Done With Blackwell Yet As GB300 & GB200 Deployments Receive Performance & Efficiency Boosting Upgrades Across The Globe
One may think that NVIDIA will move over to its Rubin GPUs now that its next-gen AI platform is being deployed around the world, but no, NVIDIA Blackwell is already deployed in countless data centers, and just like Hopper before it, the Blackwell generation continues to see optimizations, so much so that GB300 is still posting record performance in AI workloads.
According to NVIDIA, continuous optimizations to its Blackwell platform have led to significant uplift in performance and perf/watt. In just three months, NVIDIA boosted the Throughput per MW (TPS/MW) by 4x on the same GB200 NVL72 configuration running DeepSeek R1 0528 - 1K/1K.
During these four months, the company has added 38 major optimizations to the Blackwell platform after running 250,000+ simulated configurations, which add up to 1.4 million GPU Hours of optimization testing. Over 90% of these optimizations are applicable to other AI models, showcasing NVIDIA's commitment to its AI platforms even when new ones are rolling out.
On the other hand, NVIDIA's GB300 "Blackwell Ultra" platform continues to dominate AI training. On the DeepSeek-V3 671B model running on 256 GPUs, the Megatron Core reached a record-breaking 1648 TFLOPs per GPU, a 3x uplift versus GB200's 606 TFLOPs.
Pre-training DeepSeek-v3 671B with 256 GPUs achieves a world record 1,648 TFLOPs per GPU with GB300 NVL72, enabling the same training job to reach the same performance with a fraction of the hardware compared to the previous generation.
And the optimizations are highlighted once again, as the same GB300 NVL72 system now offers a 1.5x improvement versus the past 6 months.
NVIDIA also collaborates deeply with the PyTorch and JAX communities to land optimizations across its AI portfolio. As a result, using TorchTitan (PyTorch's Native Training Stack) on DeepSeek-V3 671B, the Blackwell Ultra rack delivers a 6x performance improvement versus the baseline configuration without any optimizations.
JAX sees an even bigger uplift, reaching 1025 TFLOPs/GPU and up to 4082 Tokens/s per GPU, a 10x improvement versus the optimizations from January 2026.
NVIDIA also sets up world-class scaling performance across all frameworks for pre-training, from 256 up to 1024 GPUs. For 1024 GPUs, Megatron Core holds up to 98.5% scaling while TorchTitan & JAX hold up to 97% efficiency. The main component that delivers this scaling is NVIDIA's 800 Gb/s Scale-Out networking chip within each NVL72 rack.
From Training to Inference, NVIDIA is setting some impressive records in the AI segment, and others have a lot of catching up to do, especially when NVIDIA's Vera Rubin platform is already rolling out and delivering even bigger gains over Blackwell platforms.
Follow Wccftech on Google to get more of our news coverage in your feeds.
