How the NVIDIA Vera Rubin NVL72 Delivers 10x More Tokens Per Megawatt
The thermodynamic wall of generative AI has just been pushed back. NVIDIA has officially begun ramping production of its highly anticipated NVIDIA Vera Rubin NVL72 platform, shipping rack-scale systems to its core cloud partners—including specialized GPU cloud pioneer CoreWeave, alongside hyperscalers Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure.
The 10x Benchmark: Solving the Megawatt Bottleneck
For the past three years, the generative AI race has been dominated by a single metric: raw floating-point operations per second (FLOPs). But as cluster sizes swell toward hundreds of thousands of chips, the defining constraint of the industry has shifted from silicon yields to the municipal power grid. AI labs are no longer just looking for the fastest chips; they are desperate for thermodynamic efficiency. This is where the NVIDIA Vera Rubin NVL72 makes its most disruptive statement.
According to early performance data released by CoreWeave, the Vera Rubin NVL72 architecture delivers a staggering 10x increase in throughput per megawatt compared to the previous-generation Grace Blackwell NVL72 when running the open-source reasoning model DeepSeek-R1. This is not an incremental, generational step; it is an architectural leap that fundamentally rewrites the economics of running flagship-class models.
"CoreWeave's first benchmark on DeepSeek-R1 says it all: 10x more throughput per megawatt than Grace Blackwell NVL72 — landing directly on the metric that matters most for power-constrained AI."
NVIDIA AI Blog
This 10x efficiency gain means that for every megawatt of power a data center draws from the grid, it can serve ten times as many reasoning tokens. For scale-out infrastructure providers and enterprise customers, this translates directly to a 90% reduction in power-associated operational expenditures per inference request. It turns power conservation into a competitive pricing weapon.
Architecting for the Agentic Era: The Vera CPU
At the heart of this efficiency leap is a fundamental redesign of how computing workloads are balanced. While training large language models is a highly parallelizable exercise in matrix multiplication, running advanced, multi-step "agentic" AI workflows—where models plan, execute code, call APIs, and self-correct—introduces massive serial computation overhead.
To address this, the Vera Rubin platform introduces the custom Vera CPU, built specifically to act as the traffic cop for agentic AI workloads. The Vera CPU delivers a 2x increase in single-threaded performance compared to the Grace CPU found in the Blackwell generation.
In reasoning models like DeepSeek-R1 or OpenAI's o1, the system cannot simply stream data through tensor cores. It must frequently pause to execute search algorithms, tree-of-thought logic, and structured code execution. By doubling single-threaded performance, the Vera CPU eliminates the processor bottlenecks that historically left expensive GPUs sitting idle, consuming standby power while waiting for serial instructions to complete. It is a system-level optimization that recognizes AI has transitioned from simple next-token prediction to complex computational reasoning.
The Supply Chain as an Impenetrable Moat
While competitors like AMD, Intel, and bespoke hyperscaler ASICs attempt to catch up on pure FLOPS-per-dollar, NVIDIA is scaling a different kind of barrier: physical manufacturing. The company confirmed that the Vera Rubin platform is supported by the largest, most mature rack-scale supply chain ever assembled, spanning more than 350 factory sites across 30 countries.
Building a modern AI supercomputer is no longer about plugging a PCIe card into a server. Racks like the NVL72 are incredibly complex, liquid-cooled kinetic structures that require precise plumbing, custom power distribution units, and dense NVLink switches to operate as a single logical GPU. By establishing a global, highly distributed manufacturing footprint, NVIDIA can deliver fully integrated, liquid-cooled racks directly to data center docks worldwide, bypassing the engineering bottlenecks that have delayed competitor deployments.
The Strategic Implications: Leveling the Open-Weights Playing Field
The arrival of Vera Rubin infrastructure is poised to accelerate the ongoing shift toward open-weights models. When models like DeepSeek-R1 are released open-source, the primary barrier to their widespread adoption is the cost of host infrastructure. By slashing the token cost of running DeepSeek-R1 by an order of magnitude, NVIDIA and its cloud partners are democratizing access to frontier-class reasoning.
For proprietary AI labs, this creates an aggressive pricing squeeze. If a developer can host a highly efficient, fine-tuned open model on Vera Rubin hardware for a fraction of the cost of querying a closed API, the economic gravity shifts heavily toward self-hosting on clouds like CoreWeave, Google Cloud, or Microsoft Azure. NVIDIA, in essence, is commoditizing the inference layer to ensure that regardless of which model wins, the underlying compute runs on their liquid-cooled silicon.
The Bottom Line
The NVIDIA Vera Rubin NVL72 is a clear signal that the AI infrastructure war has entered its thermodynamic phase. By prioritizing tokens-per-megawatt and optimizing the system architecture for agentic reasoning through the Vera CPU, NVIDIA has widened its lead over the industry. The bottleneck is no longer just how many chips can be printed—it is how much intelligence can be squeezed out of a single watt of electricity. On that front, Vera Rubin has set a formidable new benchmark.
This article was ultrathought.
Get breaking news, funding rounds, and analysis delivered to your inbox. Free forever.