NVIDIA Corporation is significantly reducing the memory used by its Vera Rubin NV72 rack-scale AI system in order to address the high prices and the ongoing memory shortage, according to an analysis from GF Securities. The courage followed after NVIDIA signed long term memory agreements ahead of the shortage in the memory industry which allowed it to weather the storm that most technology companies had to face due to the production constraints stemming from high demand.
Reducing Memory Capacity Can Allow NVIDIA To Cut Rubin NVL72 Rack Costs By Nearly 50%, Says Report
NVIDIA's Vera Rubin NVL72 racks are among the most powerful AI systems in the world. After being announced in January, the racks' first performance metrics surfaced earlier this month after CoreWeave shared throughput data for different AI systems. The results showed that when it came to computing mixture of expert (MoE) workloads, while systems powered by NVIDIA's Blackwell AI GPUs were able to deliver a throughput of 80,000 tokens per second at 150 megawatts of power consumption, the VR200 NVL72 Vera Rubin platform increased this by 10x to 800,000 tokens per second.
The results came after an earlier report from Bernstein outlined that an NVL72 Rubin Rack could cost $9.1 million a piece primarily due to higher memory prices. Bernstein remarked that estimates that suggested that the rack could cost $7.8 million were based on outdated estimates related to memory prices. The financial firm added that HBM4 memory prices could increase to $53/GB in 2027.
Now, a report from GF Securities suggests that NVIDIA is making adjustments to deal with the supply and price constraints in the memory market. Starting from the Vera Rubin NVL72 racks, it believes that the firm will cut the racks' SOCAMM (Small Outline Compression Attached Memory Module) capacity in half from the earlier 192GB modules to 96GB. These adjustments are in response to the constraints in the LPDDR5X market.
Similarly, the Vera CPUs will cut memory down to 28TB from the earlier 54TB to 55TB, while the GPUs will continue to use 20.7TB of HBM4 memory per rack. As for the SOCAMM for the CPU racks, the analytical firm believes that they will also adopt 96GB SOCAMM per CPU.
The financial impact of these decisions means that LPDDR5X costs for the VR200 can drop down to as low as $293,000 in case the capacity is cut down to one fourth, or drop down to $586,000 and not touch the earlier estimate of $1.2 million. GF Securities adds that without the adjustments, memory costs could sit at 29% of the total bill of material costs of the VR200 system at $2.1 million. The 29% is higher than the preferred level of 20%.
Follow Wccftech on Google to get more of our news coverage in your feeds.
