NVIDIA’s Vera Rubin NVL72 Claims 30x Efficiency Leap Over GB300 NVL72 for AI Agents

NVIDIA's Vera Rubin NVL72 Claims 30x Efficiency Leap Over GB300 NVL72 for AI Agents

First on-silicon benchmarks show NVIDIA’s Vera Rubin NVL72 system delivering up to 35x lower token costs for agentic AI workloads compared with the GB300 NVL72.

There’s a new benchmark in AI hardware, and NVIDIA wants you to know about it. On 24 August 2026, the company posted first on-silicon performance results for its Vera Rubin NVL72 system — and the numbers it’s putting forward are striking enough to get the data-centre industry talking.

The headline claim: Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than the previous-generation GB300 NVL72 when running agentic AI workloads. Alongside that, NVIDIA says the new system achieves up to 35x lower cost per million tokens on the same tasks. Both figures come from NVIDIA’s own testing using the SemiAnalysis AgentX benchmark, running the DeepSeek V4 Pro model on what the company describes as real-world “agentic coding trajectories.”

NVIDIA was direct about the context. The @nvidia account posted that these results apply specifically to agentic sessions — not chat or summarisation workloads. That distinction matters, because the two types of task put AI hardware under very different kinds of pressure.

What Makes an Agentic Workload Different?

A standard chatbot interaction is relatively short. You ask a question, the model responds, and that’s broadly it. An AI agent is something else entirely. It might spend minutes or hours working through a problem — writing and testing code, calling external tools, planning a sequence of actions, and keeping track of everything it’s done along the way. The context window — the amount of information the model holds in memory at once — grows continuously throughout that process.

That’s a fundamentally different demand on the hardware. Energy efficiency per unit of work, and cost per token, become the metrics that actually matter for anyone running agents at scale.

NVIDIA has designed Vera Rubin NVL72 with exactly that in mind. The system follows on from GB300 NVL72, which itself was no slouch — a rack-scale platform combining 72 Blackwell Ultra GPUs and 36 Arm-based Grace CPUs in a single NVLink-connected domain, drawing somewhere around 120–150 kW under full load, with roughly 20 terabytes of pooled HBM3e memory and aggregate FP4 performance in the 1.0–1.1 exaFLOPS range.

And GB300 NVL72 was already a big step up from NVIDIA’s Hopper-based systems. Official NVIDIA documentation puts it at 10x higher tokens per second per user and 5x higher tokens per second per megawatt than Hopper, adding up to a claimed 50x improvement in overall AI factory output. Vera Rubin NVL72, if the benchmarks hold, takes that further still.

The Numbers — and the Caveats

Thirty times more throughput per megawatt. Thirty-five times lower cost per million tokens. Those are the kind of figures that make cloud operators and data-centre managers pay close attention.

But independent analysts have been quick to add context. The 30x and 35x figures are vendor-reported results from a specific benchmark setup — SemiAnalysis AgentX with DeepSeek V4 Pro — and haven’t been independently audited by any official standards body. No UK regulator, including Ofgem or the Department for Energy Security and Net Zero, currently publishes benchmarks for AI systems tokens per second per megawatt. These metrics are, for now, largely vendor-defined.

It’s also worth being clear that benchmark performance on one workload doesn’t automatically translate across all tasks. Agentic coding is what was tested here. Multimodal models, fine-tuning jobs, or non-agentic inference could tell a different story.

There’s another consideration that environmental groups and some analysts raise: the rebound effect. When the cost per unit of AI work falls sharply, organisations tend to run more of it. Higher efficiency per transaction doesn’t always mean lower total energy use — it can mean the same power budget gets used to run far more workloads. Whether Vera Rubin NVL72 ultimately reduces data-centre energy consumption or simply enables a larger volume of AI tasks within the same footprint is a question that real-world deployments will have to answer.

What NVIDIA Says It’s Built

The name Vera Rubin follows NVIDIA’s tradition of naming platform generations after scientists. Vera Rubin was a pioneering astronomer whose work on galaxy rotation curves provided some of the strongest early evidence for dark matter. The choice of name signals that NVIDIA sees this as a generational step, not an incremental update.

Jensen Huang, NVIDIA’s chief executive, has previously argued that AI agents represent the next major wave of enterprise computing — systems that don’t just answer questions but carry out extended tasks autonomously. The Vera Rubin NVL72 announcement is NVIDIA’s clearest signal yet that it’s building hardware specifically optimised for that workload class, rather than treating agents as simply a longer version of standard inference.

The company argues that the efficiency gains make it economically viable to run AI agents “continuously, at scale” — a claim that, if borne out in production, could shift how cloud providers price and provision agentic services.

What Happens Next

No specific deployment numbers for Vera Rubin NVL72 — globally or in the UK — have been publicly disclosed. Cloud providers and data-centre operators will be evaluating the system, and broader availability timelines will depend on those partners. Independent benchmarking, when it comes, will be the real test of whether NVIDIA’s on-silicon results hold up across a wider range of workloads and configurations.

What This Means for Kent Residents

Kent sits within a corridor of data-centre development serving London and the South East, and if UK cloud providers adopt Vera Rubin NVL72 in their facilities, the efficiency gains could allow more AI services to be delivered within the same power envelope — which matters for grid operators like UK Power Networks that manage electricity distribution across the county. For residents using AI-powered services through the NHS, banks, or government platforms, any hardware improvements tend to be invisible at the point of use, but could over time translate into faster, more reliable services if providers pass the benefits on. Whether that happens — and at what pace — will depend entirely on commercial decisions made well above the level of the average Kent household.

Source: @nvidia

NVIDIA's Vera Rubin NVL72 Claims 30x Efficiency Leap Over GB300 NVL72 for AI Agents Quiz

5 questions