Nvidia’s Rubin platform starts shipping as AI data centres race for more compute
Nvidia's next-generation Rubin lineup — six chips built around a new GPU, CPU and networking stack — is beginning to reach data centre customers this half, promising sharply lower per-token costs for the largest AI models.
Nvidia’s next-generation AI computing platform, Rubin, is beginning to ship to data centre customers this half of the year, moving the company’s roadmap forward roughly a year after Blackwell became the industry’s default training and inference engine.
Rubin is not a single chip but a full rack-scale system: a new Rubin GPU paired with a “Vera” CPU, a sixth-generation NVLink switch, updated networking silicon, and a new rack-scale configuration Nvidia calls Vera Rubin NVL72. The company says the platform can cut the cost of generating each token of AI output by an order of magnitude compared with Blackwell, and needs meaningfully fewer GPUs to train large mixture-of-experts models — the architecture increasingly favoured by frontier labs because it activates only part of a model’s parameters for any given request.
For an industry where compute has become the primary bottleneck on how quickly new models can be trained and served, the arrival of a materially more efficient platform matters beyond Nvidia’s own balance sheet: cloud providers, sovereign AI programmes and the major model developers all size their roadmaps around what hardware will be available and at what cost per unit of intelligence. Executives at several frontier labs have publicly tied their own scaling plans to Rubin’s availability, underscoring how concentrated the AI compute supply chain remains around a single vendor’s release cadence.
