NVIDIA’s new rack-scale AI system reportedly delivers ten times the energy efficiency of its Blackwell predecessor on a key reasoning workload, according to partner benchmark data published in late July 2026.
NVIDIA has publicised a striking efficiency claim for its Vera Rubin NVL72 system: ten times more AI tokens per megawatt than the GB200 NVL72 Blackwell rack it replaces, based on tests run by cloud provider CoreWeave on the DeepSeek-R1 reasoning model.
The figures appeared on NVIDIA’s official social media and product pages around 21 July 2026, billed as “first measured silicon” performance data for Vera Rubin NVL72. CoreWeave ran DeepSeek-R1 on both systems using NVIDIA TensorRT-LLM and NVIDIA Dynamo software, holding user responsiveness constant — measured as tokens per second per user — to make the comparison fair.
What the Benchmark Actually Shows
The headline metric is tokens per second per megawatt — a measure of how much AI output a system squeezes from each unit of electrical power. At around 150 tokens per second per user, CoreWeave’s chart shows Vera Rubin NVL72 pulling ten times the throughput per megawatt compared with GB200 NVL72.
That’s a big number. But it’s a specific number, tied to one model, one operating point, and one benchmark configuration.
NVIDIA’s own product materials state the system can deliver “up to 10x more tokens per megawatt than GB200 NVL72” and “one-tenth the cost per million tokens” on DeepSeek-R1 under those specified conditions. The company frames Vera Rubin NVL72 as a rack-scale, agentic AI supercomputer built to maximise performance per watt — the successor to GB200 NVL72 in its Oberon architecture line.
The Caveats Analysts Are Flagging
Not everyone is ready to take the 10x claim at face value. As of late July 2026, there is no MLPerf submission, peer-reviewed paper, or independently audited benchmark publicly confirming the figure. The data comes from CoreWeave — which is both an NVIDIA cloud partner and an NVIDIA investee — making this first-party or partner evidence rather than independent testing.
Technical commentary from outlets including Semianalysis, TECHi, and The Agent Times points out that the benchmark covers a single reasoning model at a single operating point. Real-world gains will depend on workload mix, use, data centre design, and software stack. The 10x figure does not automatically mean 10x lower electricity bills, 10x lower total AI service costs, or 10x performance across all AI models.
Analysts at the Futurum Group also note that Vera Rubin NVL72’s firmware and software kernels are at an early stage of maturity compared with Blackwell. Performance and efficiency may shift as software optimisations develop.
Why Energy Efficiency Has Become the Central Battleground
The focus on tokens per megawatt isn’t accidental. Power-supply constraints have become one of the biggest bottlenecks in AI data centre expansion. Grid capacity limits how much compute operators can physically run, so doing more AI work within a fixed power envelope has real commercial value.
NVIDIA positions Vera Rubin NVL72 squarely in that context, targeting what it calls “agentic AI” workloads — models like DeepSeek-R1 that perform complex reasoning and multi-step tasks. These workloads are driving current growth in AI inference demand, and they’re power-hungry.
AI developers and enterprises will likely want more transparent, independently verified benchmarks before making large procurement or migration decisions on the back of the 10x claim alone. That’s a reasonable position given the source.
What NVIDIA and CoreWeave Are Saying
NVIDIA’s materials are unambiguous about the headline figure. CoreWeave, for its part, describes the results as live hardware numbers from real silicon, not modelled projections. Joe Pierce, an industry figure who shared a summary of the benchmark on LinkedIn, highlighted the tokens-per-second-per-megawatt chart as evidence of a step change in inference efficiency for reasoning models.
The company has not published the full methodology or made the raw benchmark data available for independent reproduction — a gap that critics say limits confidence in the broader claim.
There is currently no UK public body — not Ofgem, not the Department for Energy Security and Net Zero, not the Office for National Statistics — that has independently assessed or validated the Vera Rubin NVL72 performance figures. All reported data originates from NVIDIA and CoreWeave.
What This Means for Kent Residents
For most Kent residents, the immediate impact is indirect. If efficiency gains of this scale hold up under independent scrutiny, they could reduce the power demand of AI data centres over time — relevant in a county where planning applications for large compute facilities raise questions about grid capacity, land use, and local environmental impact. Bodies such as Kent County Council and NHS Kent and Medway ICB rely on cloud-based AI tools, and more efficient infrastructure at major providers could eventually affect the cost and performance of those digital services, though any concrete benefit to local users has not been quantified. For now, the 10x claim remains vendor-reported data, and Kent residents — like everyone else — are waiting on independent verification before drawing firm conclusions.
Source: @nvidia
NVIDIA Vera Rubin NVL72 Claims 10x More AI Tokens Per Megawatt Than Blackwell in CoreWeave Benchmark Quiz
5 questions