Jul 22, 2026
Enterprise

Nvidia positions Vera Rubin as an AI factory CPU and networking play

Nvidia says Vera Rubin combines custom Arm CPUs, GPUs, networking and DPUs to cut token costs and improve performance per watt for agentic AI workloads.

Dominic Okoye

By Dominic Okoye · Staff Writer

· 3 min read

Nvidia positions Vera Rubin as an AI factory CPU and networking play
Photo: SiliconANGLE

Nvidia has launched Vera Rubin, a full-stack AI infrastructure platform with pricing and revenue impact not disclosed, built around seven co-designed chips and a broader rack-to-cluster architecture. The company says the platform is intended to improve performance per watt and reduce token costs, while giving Nvidia more control over the CPU, networking and data processing layers around its GPUs.

Nvidia said production is ramping with cloud and AI infrastructure partners including CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. The company did not disclose deployment volumes, customer spend, or comparative token-cost figures for those early rollouts.

The strategic claim is straightforward: Nvidia says agentic AI workloads put the CPU back into the inference path. In a briefing, Nvidia’s Hannah Coutand said “AI is asking for a new CPU” because agentic systems require processors to handle tool calls, code execution, queries, orchestration and data movement between GPU reasoning steps.

Vera’s CPU target

The Vera CPU is based on Arm and uses Nvidia’s custom Olympus core. Nvidia disclosed 88 custom cores, 176 hardware threads, up to 1.2 terabits per second of LPDDR5X memory bandwidth, 164 megabytes of unified L3 cache and up to 1.8 terabytes per second of coherent CPU-to-GPU bandwidth through NVLink-C2C.

Nvidia describes the design goal as high single-threaded CPU performance at scale. Coutand said Nvidia “didn’t set out to go win CPUs,” and framed the chip as a response to the CPU becoming a bottleneck in what the company calls the AI factory.

Nvidia’s Ian Finder said Olympus was designed with a 10-wide decode engine, aggressive reordering logic and a graph prefetcher aimed at pointer-heavy workloads such as compilers, graph structures and agent runtimes. The company’s pitch is less about replacing every general-purpose data center CPU and more about owning the processor role inside GPU-heavy AI systems.

Networking becomes part of the system sale

Vera also extends Nvidia’s argument that AI infrastructure is sold as a system, not as a rack of interchangeable components. Inside the CPU, Nvidia’s second-generation Scalable Coherency Fabric links cores, cache, LPDDR5X controllers, I/O and NVLink-C2C interfaces. Nvidia contrasts that approach with chiplet-based CPUs, which it says can incur higher latency and lower effective bandwidth when traffic crosses die boundaries.

The platform also leans on Spectrum-X for scale-out networking. Nvidia’s stack includes 102.4T Spectrum-6 switches, 1.6T ConnectX-9 SuperNICs, adaptive routing, congestion control, telemetry and software tuned for RDMA and AI traffic patterns. NVLink handles scale-up connectivity inside a rack, while Spectrum-X is positioned for connecting racks across larger AI clusters.

That matters commercially because Nvidia is using Rubin-class systems to pull adjacent silicon into deals where GPUs already control the budget. The platform includes Vera Rubin NVL72, Vera CPU racks, BlueField-4 infrastructure processors and Spectrum-6 switching, all pitched as a co-designed system.

ZK Research analyst Zeus Kerravala argued that CPUs may be Nvidia’s next share-gain opportunity, especially in AI clouds, hyperscalers and model builders running agentic or reinforcement-learning-heavy workloads. AMD is holding its own Advancing AI summit this week, which should show how it plans to answer Nvidia’s CPU and networking claims.

This story draws on original reporting from SiliconANGLE.

More from Enterprise

All Enterprise →