Jul 24, 2026
Enterprise

AMD AI PCs pitched as a way to cut agentic AI cloud costs

AMD says Ryzen AI Halo can run agentic AI workloads locally as CIOs question cloud token costs, security and latency.

Dominic Okoye

By Dominic Okoye · Staff Writer

· 3 min read

AMD AI PCs pitched as a way to cut agentic AI cloud costs
Photo: SiliconANGLE

AMD AI PCs are being positioned as a local compute option for enterprises trying to keep agentic AI from becoming a larger cloud bill. Rahul Tikoo, senior vice president and general manager of AMD’s Client Business Unit, said at AMD Advancing AI 2026 that CIOs are weighing whether some AI agent work should move from cloud infrastructure to endpoint devices.

The argument is straightforward: as AI moves from chatbot-style assistance toward autonomous agents, usage can multiply and token-based cloud charges can become a material operating cost. Tikoo said enterprises still expect to use cloud and private infrastructure, but are also looking at AI PCs as part of a hybrid setup where some inference happens closer to users and data.

AMD’s product answer is Ryzen AI Halo, a platform the company says supports local AI model execution and agentic workflows on a single system. The platform can be configured with up to 128GB of memory, according to AMD. The company did not disclose pricing, customer adoption figures or deployment volumes in the discussion.

How could AMD AI PCs reduce agentic AI token costs?

Cloud AI services often charge based on tokens, the units of text processed by a model. Running some inference locally can reduce the amount of work sent to paid cloud models, though AMD did not quantify the savings or provide enterprise benchmarks.

Tikoo said one use case is giving developers a local environment for building and testing AI applications without paying for every experiment through cloud token consumption. He also argued that endpoints can be a practical place to run inference because user and enterprise data often already sits there, which may also help address latency and security requirements.

AMD’s claim depends partly on the improving capability of smaller models. Tikoo said models with 9 billion or 24 billion parameters are reaching quality levels that make them usable for tasks that previously required frontier systems. That is a company executive’s assessment, and the interview did not include third-party model evaluations or named enterprise deployments.

The enterprise pitch also reflects AMD’s broader silicon position. Tikoo said AMD can supply compute from cloud and data center systems to endpoints and edge devices, and that the company is working on software to make those layers easier to connect. For buyers, that software layer will be the harder test. Hardware that can run local models is useful only if IT teams can manage policies, data access, model updates and security across fleets of PCs.

The remarks came during an interview with theCUBE hosts John Furrier and Dave Vellante at AMD Advancing AI 2026. TheCUBE disclosed that it was a paid media partner for the AMD event and said AMD did not have editorial control over the coverage.

This story draws on original reporting from SiliconANGLE.

More from Enterprise

All Enterprise →