Skip to content
Eton Technology SolutionsEton Technology Solutions

Vera Rubin NVL72 claims 30x more work per watt: why efficiency is now the AI infrastructure metric that counts

September 02, 2026

Vera Rubin NVL72 claims 30x more work per watt: why efficiency is now the AI infrastructure metric that counts

NVIDIA has published measured data putting Vera Rubin NVL72 at up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads. For power-limited sites, that number matters more than raw performance.

What was announced

  • NVIDIA reports Vera Rubin NVL72 delivers up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads, measured on the DeepSeek V4 Pro model.
  • NVIDIA also reports up to 35x lower cost per million tokens versus GB300 NVL72.
  • The measurements used the SemiAnalysis AgentX workload of recorded real-world agentic coding sessions; NVIDIA states the results are pending SemiAnalysis review and do not yet reflect Vera CPU performance for tool calling.
  • NVIDIA cites OpenRouter data indicating agentic AI workloads consume 15x more tokens than a simple chat request.
  • NVIDIA reports GB300 NVL72 delivers up to 15x better throughput per megawatt than the Hopper architecture on the same model.
  • NVIDIA states its DSX MaxLPS power management can provision up to 40% more GPUs within the same megawatt budget.

The ETON view

These are vendor-measured figures on a vendor-selected workload and they are pending third-party review, so treat the multiple as directional rather than as a specification. The underlying shift is real though, and it is the one worth acting on: the industry has stopped quoting performance and started quoting performance per megawatt, because power is what is actually scarce.

That reframes a lot of British infrastructure planning. If your constraint is a fixed power allocation in a colo or a building you cannot easily upgrade, the right question is how much useful work you can extract per kilowatt, and the answer usually involves cooling and rack design as much as silicon. It also cuts the other way for anyone running inference at modest scale. A 15x or 30x efficiency claim against a flagship rack system says little about a four-GPU inference box in a 6 kW rack, where the sensible move is often current-generation hardware bought well rather than next-generation hardware bought early. We size to the power you have, not to the press release.

Related infrastructure

Category: AI & GPU · Vendor: NVIDIA · Technology: Vera Rubin NVL72, GB300 NVL72, NVFP4, agentic inference · Last verified: 02 Sep 2026 · ~2 min read

Sourcing this kind of infrastructure? Talk to ETON about availability, lead time and pricing.

Cart 0

Your cart is currently empty.

Start Shopping
Call Us