Jul 24, 2026
Enterprise

AMD token routing push aims at enterprise AI cost overruns

AMD says routing AI workloads to MI350P GPUs cut its token bill 43%, as enterprises rethink cloud-heavy AI deployments.

Colin Brandt

By Colin Brandt · Enterprise Reporter

· 3 min read

AMD token routing push aims at enterprise AI cost overruns
Photo: SiliconANGLE

AMD token routing is becoming a cost argument as the chipmaker courts enterprises that have moved AI projects past experiments and into production. John Hampton, corporate vice president of global enterprise technical sales at Advanced Micro Devices Inc., said AMD’s own pilot routing workloads to MI350P GPUs reduced its token bill by 43% and improved response speed by 2.9x.

Hampton discussed the results during an interview with theCUBE hosts Dave Vellante and John Furrier at AMD Advancing AI 2026. The company did not disclose the pilot’s absolute spending level, workload mix, duration or scale, which limits how much buyers can read across to their own environments.

The pitch is aimed at a familiar problem for enterprise IT teams: early AI infrastructure choices are now showing up as recurring operating costs. Hampton said many companies began with large GPU clusters from AMD rivals and ran broad workloads on cloud-based frontier models, then reassessed after seeing the bill. He did not name the competing vendors.

What is AMD token routing?

AMD’s framing of token routing means sending each AI request to the lowest-cost infrastructure that can handle the job, rather than defaulting to a frontier model running on a large GPU cluster. In some cases, Hampton said, inference can run on CPUs or lower-cost, lower-power GPUs instead.

That distinction matters as enterprise AI shifts from employee prompting toward more automated workflows. More autonomous systems can generate far more model calls, which pushes token consumption from a line item for pilots into a budget issue for production operations.

Hampton described token economics as the top issue he is hearing from enterprise customers, saying affordability and return on investment have come under pressure. The company’s answer is a hybrid operating model that places workloads across CPUs, lower-cost GPUs and higher-end accelerators depending on performance and cost requirements.

AMD’s internal example used MI350P GPUs in place of frontier cloud models for selected workloads. AMD presented the 43% bill reduction and 2.9x response-speed increase as evidence that routing decisions can improve both cost and performance, though the company did not publish a technical breakdown of the benchmark or identify the model classes involved.

Why enterprises are revisiting AI infrastructure now

The cost discussion is tied to broader data center planning. Hampton said companies that consolidate data centers can free up power and physical space, then use that capacity for AI infrastructure. That makes token routing less a standalone optimization and more part of a larger infrastructure refresh decision.

The message also gives AMD a way to compete for enterprise AI workloads without arguing that every deployment needs the largest possible accelerator cluster. For buyers, the practical question is whether routing policies, model selection and hardware placement can be governed tightly enough to reduce spend without creating latency, accuracy or operational problems elsewhere.

SiliconANGLE Media said theCUBE was a paid media partner for the AMD Advancing AI event and that AMD and other sponsors did not have editorial control over the coverage.

This story draws on original reporting from SiliconANGLE.

More from Enterprise

All Enterprise →