Jul 31, 2026
Enterprise

Physical AI infrastructure strain grows as robotics moves beyond cloud

Executives at theCUBE’s robotics event said robots and edge AI are shifting demand toward cheaper inference, secure GPUs and local models.

Dominic Okoye

By Dominic Okoye · Staff Writer

· 4 min read

Physical AI infrastructure strain grows as robotics moves beyond cloud
Photo: SiliconANGLE

Physical AI infrastructure is becoming a more specific problem as robots, autonomous systems and edge devices move AI workloads outside standard cloud patterns, executives said during SiliconANGLE’s theCUBE + NYSE Wired: Robotics & AI Infra Leaders event. No companies disclosed funding, revenue, customer counts or new commercial terms, but several vendors described how they are positioning around inference cost, GPU access, security and local model deployment.

Haseeb Budhani, co-founder and chief executive of Rafay Systems, said the company is using orchestration to let providers run open models without assigning whole GPU systems to a single customer. He argued that confidential GPU slicing can improve security while helping providers sell more of their infrastructure time and giving enterprises a lower-cost option than dedicated capacity.

The claim fits a wider shift in AI infrastructure sales. As enterprises move from experiments to workloads tied to factories, vehicles and distributed devices, the buying discussion is less about model demos and more about utilization, isolation and predictable economics. Rafay did not provide pricing, margin impact or customer deployment data.

What is physical AI infrastructure?

Physical AI infrastructure refers to the compute, networking, data access, security and device-side systems needed to run AI in robots, machines, vehicles and other real-world environments. Unlike centralized chatbot workloads, these systems often need low-latency inference, local processing and controls that work when connectivity or power is constrained.

Max Kan, tokenomics technical lead at SemiAnalysis, said agentic AI applications consume more tokens than basic chat because each action or follow-up can require earlier context to be processed again. He said that makes token volume, model choice and workload design central to infrastructure planning.

Positron AI is targeting inference economics with systems optimized around tokens per dollar and tokens per watt, according to Darren Chien, the company’s managing director for APAC. Chien said inference should be treated as an economics problem, including for enterprise data centers that cannot take on liquid-cooled racks and need air-cooled systems.

Axiado is approaching the problem through silicon-level platform management. Founder, president and Chief Executive Gopi Sirineni said the company’s systems monitor elements such as fans and liquid cooling while adjusting frequency and voltage by workload, offloading some management work from CPUs and GPUs.

Edge AI pushes models and networking closer to devices

Several executives argued that smaller, customizable models will be part of the answer for edge AI. Ramin Hasani, co-founder and chief executive of Liquid AI, said compact models can be customized for specific tasks on laptops, vehicles and other devices, lowering inference costs and reducing reliance on centralized data centers. He also said a single small model should not be treated as generally intelligent.

Networking is also being pulled into the application layer. Mansour Karam, founder and chief executive of Aria Networks, said operators need high-resolution telemetry, including microsecond-level data, to spot failures and performance problems before they affect expensive AI workloads. He framed the company’s role as helping human operators make faster decisions rather than removing them from network operations.

H Company is focusing on computer-use agents for enterprises, according to Chief Executive Gautier Cloix. He said the company designs agents so they cannot complete irreversible actions, such as sending an email, deleting information or placing an order, without human approval at first.

GPU capacity is becoming a market of its own

Compute availability and pricing are creating room for financial and marketplace products around GPU capacity. Carmen Li, founder and chief executive of Silicon Data and chief executive of The Compute Exchange, said her companies are developing compute benchmarks and instruments intended to help infrastructure operators and large AI buyers manage exposure to GPU price swings.

San Francisco Compute Co. is building a marketplace for GPU capacity, according to Chief Business Officer Alan Butler. He said the company matches workloads with infrastructure based on location and availability requirements and backs listings with service-level agreements from data center providers.

Axelera AI is extending its edge-focused architecture into servers and cloud systems, said Alexis Crowell, chief marketing officer and general manager of the Americas. She said the company is using familiar software frameworks so developers can adopt its hardware without rewriting applications around proprietary tooling.

This story draws on original reporting from SiliconANGLE.

More from Enterprise

All Enterprise →