NVIDIA confirms fully automated manufacturing for its Vera Rubin NVL72 AI racks, as Microsoft becomes the first hyperscale cloud provider to power on and validate the new system.
NVIDIA has confirmed that production of its Vera Rubin NVL72 rack-scale AI system is now under way, with Microsoft becoming the first hyperscale cloud provider to power on and validate the new hardware in its facilities.
The announcement marks a major step forward for NVIDIA’s Rubin generation platform — the successor to its Grace Blackwell systems — and signals that the most powerful AI infrastructure yet built is moving from engineering demonstration into real-world deployment.
One Minute, Zero Cables, No Fans
The headline engineering claim is striking in its simplicity. Each Vera Rubin NVL72 compute tray is assembled in around one minute using fully automated manufacturing — a process NVIDIA says involves no human hands on the line.
That compares to assembly times of roughly 90 to 120 minutes reported for certain earlier Blackwell and GB200 compute trays, though exact figures vary by system and source. All accounts agree the reduction is dramatic.
The tray itself is cable-free, hose-free, and fan-free. There are no data cables snaking across the board, no air-cooling fans, and no flexible hoses to connect. Instead, the entire system runs on 100% liquid cooling, with coolant circulating at around 45°C. NVIDIA says this design makes automated assembly far more reliable and repeatable than earlier generations.
Each NVL72 rack holds 72 Rubin GPUs and 36 Vera CPUs, all interconnected via sixth-generation NVLink with an aggregate bandwidth of around 260 terabytes per second. Industry reporting, based on NVIDIA vendor briefings, puts AI inference performance at multi-exaflop levels using low-precision formats such as FP4 — though specific figures like 3.6 exaflops FP4 have not yet been directly confirmed in publicly available NVIDIA documentation.
Microsoft Gets There First
NVIDIA publicly congratulated Microsoft for hosting the first operational Vera Rubin NVL72 system. The rack is currently running in Microsoft’s labs for validation — meaning it’s powered on and working, but not yet serving customer workloads through Azure at scale.
Microsoft has said it plans to roll Vera Rubin NVL72 out across Azure data centres over the coming months. The company frames the early deployment as giving Azure a first-mover advantage in next-generation AI infrastructure, continuing a pattern of large-scale NVIDIA GPU deployments that stretches back through Grace Blackwell and earlier generations.
But industry analysts have urged some caution about the word “operational.” Being first to power on a system for lab validation is not the same as making it generally available to customers. The gap between delivery, validation, and production deployment can run to many months for hardware this complex.
Infrastructure That Demands New Data Centres
The NVL72 rack isn’t something you slot into an existing server room. It runs on an 800-volt DC power architecture and requires full liquid-cooling infrastructure throughout. Traditional air-cooled data halls simply can’t host it.
That requirement has cost implications too. Some industry sources put the price of a single NVL72 rack at somewhere in the region of USD 7 to 8 million — roughly £5.5 to £6.3 million at current rates — though NVIDIA and Microsoft have not confirmed any official pricing, and those figures should be treated as unverified estimates.
The Rubin platform also integrates with NVIDIA’s broader ecosystem, including ConnectX-9 SuperNICs and BlueField-4 data processing units, which handle networking and offload infrastructure tasks from the main compute.
The Bigger Picture
Vera Rubin NVL72 is NVIDIA’s third generation of what the company calls rack-scale co-design — an approach where the rack, trays, cooling, power delivery, and networking are engineered together from the start rather than bolted together from separate components. NVIDIA says this drives down token cost and improves performance per watt for large-scale AI workloads including generative AI, agentic AI, and large language model serving.
Satya Nadella, Chief Executive of Microsoft, has not made a separate public statement on the Vera Rubin NVL72 deployment beyond what has been shared through NVIDIA’s announcement and Microsoft’s own communications about Azure AI infrastructure.
Critics have raised questions about energy use. Racks this dense, deployed at scale, carry significant power and cooling demands. Environmental groups and grid planners in several countries are already watching hyperscale AI infrastructure expansion closely, given its potential impact on local electricity grids and carbon targets.
The move to fully automated manufacturing reflects a broader shift in high-end server production. Robotic assembly lines handle complex, high-density components more consistently than manual labour, and at the volumes hyperscale cloud providers now require, the speed advantage compounds quickly.
What This Means for Kent Residents
There’s no public evidence that Vera Rubin NVL72 racks are heading to data centres in Kent specifically — NVIDIA and Microsoft haven’t disclosed site-level detail for UK deployments. For most Kent residents, the practical effect will be indirect: as these systems move from lab validation into Azure production over the coming months, AI-powered tools and cloud services used by local businesses, schools, and public bodies such as NHS Kent and Medway ICB could eventually run faster or handle more complex tasks. The energy and planning implications of large-scale liquid-cooled AI infrastructure in the UK are a matter for Ofgem and National Grid ESO rather than local councils for now, but that picture could change if hyperscale providers announce new UK data centre sites.
Source: @nvidia
NVIDIA Vera Rubin NVL72 AI Racks Enter Automated Production as Microsoft Powers On First System Quiz
5 questions