Nvidia posts Vera Rubin benchmark claims for AI data centers
Nvidia says its next AI platform delivers large efficiency gains, with CoreWeave reporting 10 times more tokens per watt on DeepSeek R1.
By Wei-Lin Zhao · AI Correspondent
· 3 min read
Nvidia released new performance data for its Vera Rubin platform, claiming large gains in efficiency, CPU performance, cooling and networking as it prepares the system for broader customer deployment. The numbers matter because Nvidia is selling more than chips: it is pitching a full-stack design for so-called AI factories at a time when power, cooling and networking are constraining AI buildouts.
CoreWeave, the GPU cloud provider and Nvidia partner, reported that early Vera Rubin production runs delivered 10 times more tokens per watt on DeepSeek’s R1 model than Nvidia’s prior GB200 NVL72 system based on Blackwell. Nvidia also said software and system tuning applied to the GB200 NVL72 lifted throughput per megawatt by more than four times over three months.
Nvidia said those efficiency gains were validated across more than 250,000 configurations and over 1.4 million GPU hours of testing. The company did not disclose pricing for Vera Rubin systems, deployment volumes or a firm general availability date in the information released.
Vera CPU targets agent workloads
The company also published benchmark claims for its Vera central processing units, which are intended to support autonomous AI agents that run less predictable workloads than conventional inference jobs. Nvidia said the Vera CPUs are based on a custom microarchitecture specification called Olympus core and are tuned for the irregular control flows associated with agentic AI.
In Nvidia’s tests, Vera delivered 1.9 times faster agentic performance and a sixfold latency improvement compared with x86 alternatives. The company also said Vera beat AMD’s flagship EPYC Turin CPU by nearly 100% on selected industry benchmarks. The “selected” qualifier matters: Nvidia is making the case for vertical integration, but the released figures do not amount to an independent across-the-board CPU comparison.
Power, cooling and network claims
Nvidia framed the Vera Rubin update as a data center architecture story rather than a single accelerator refresh. The company said optimizing infrastructure and energy systems dynamically would allow operators to place 40% more GPUs inside the same power envelope.
For cooling, Nvidia pointed to a 45-degree Celsius closed-loop liquid-cooling setup that it said can save about 4 million gallons of water per megawatt each year compared with standard cooling methods. That claim speaks directly to one of the harder constraints around AI data center expansion, though Nvidia did not provide site-level deployment data alongside the figure.
Networking remains part of the pitch. Nvidia said its sixth-generation NVLink 6 interconnect produced 2.3 times higher simulated decode throughput for large language models than Ethernet-based networks. Its Spectrum-X platform, according to the company, delivered 1.6 times faster remote direct memory access bandwidth while using 1.7 times fewer switches. Nvidia also claimed five times better optical power efficiency and 10 times higher reliability for the integrated network.
The company said Spectrum-6, the latest Spectrum-X generation, is arriving at AI factory deployments for CoreWeave, Microsoft, Nebius, SpaceXAI and Tesla.
Large buyers are in line
Nvidia said it is working to ship Vera Rubin systems to customers and partners including Google Cloud, Microsoft Azure, Meta, Oracle Cloud Infrastructure, Dell Technologies, OpenAI and CoreWeave. That list shows the intended market clearly: hyperscalers, model companies and GPU cloud providers with enough capital and power access to absorb full-rack AI systems.
The update also reinforces Nvidia’s current strategic position. Competitors can contest individual chips, CPUs or networking components, but Nvidia is trying to make the buying decision about an integrated compute fabric. The benchmarks are vendor claims and partner-reported results, but they show where Nvidia wants the market to focus: tokens per watt, throughput per megawatt and fewer bottlenecks across the full AI data center stack.
This story draws on original reporting from SiliconANGLE.