AMD AI token economics pitch targets enterprise inference costs
AMD says enterprise AI buyers are shifting from pilots to outcome-based deployments, putting inference costs and hybrid model choices in focus.
By Colin Brandt · Enterprise Reporter
· 3 min read
AMD AI token economics is becoming a central part of the chipmaker’s enterprise pitch as companies move AI projects from experiments into production, according to Suresh Andani, AMD’s corporate vice president of compute and enterprise AI. AMD did not announce pricing, a customer contract or new revenue figures in the remarks, but Andani argued that infrastructure decisions are increasingly being judged by inference cost, deployment flexibility and business outcomes rather than hardware specifications alone.
Andani made the comments during an interview with SiliconANGLE’s theCUBE at the AMD Advancing AI 2026 event. TheCUBE disclosed that it was a paid media partner for the event and said AMD did not have editorial control over its coverage.
The shift AMD is describing is familiar to enterprise technology vendors: AI budgets are moving from proof-of-concept spending toward systems that need to justify themselves in production. In that setting, the model choice matters because token consumption can become a recurring cost center, especially when teams route tasks to frontier models by default.
What is AMD saying about AI token economics?
Andani said enterprises are learning that many AI tasks do not require frontier models, even though those models remain useful for workloads needing deeper context or high concurrency. In AMD’s view, the cost per useful AI output is pushing buyers toward a mix of cloud-based frontier APIs and open-weight models hosted on infrastructure they control.
That framing puts AMD’s enterprise AI strategy around hybrid deployment in two senses. Workloads may run in the cloud, on-premises or at the edge, while companies may also choose between frontier models and open-weight models depending on the job. Andani said enterprises are looking at open-weight models on-premises for cost control, security, data sovereignty and operational control.
AMD’s position is that it wants to support both sides of that setup: access to frontier APIs through cloud channels and infrastructure for hosting open-weight models inside enterprise environments. Andani said the topic comes up in enterprise conversations, though AMD did not provide customer names, deployment counts or comparative cost data in the interview.
Outcome-based AI buying changes the infrastructure sale
Andani said enterprise customers are buying against business outcomes rather than buying chips or servers as standalone assets. He gave the example of an oil and gas company using AI to reduce contract leakage, where the measure of value is whether the system cuts lost dollars from contracts.
That does not make the infrastructure layer irrelevant. Andani said total cost of ownership is still affected by the foundational technology, including compute performance and system design. His point was that “speeds and feeds” now need to connect back to an enterprise outcome, rather than sit as a separate hardware benchmark.
The comments also show AMD trying to broaden the conversation beyond GPUs. Andani described a stack that can include GPUs, CPUs, networking chips, platforms, an AI control plane and independent software vendor applications. Some enterprises want to assemble their own AI agent stacks using platforms from OEM and cloud partners, he said, while others want more complete AI platforms and prebuilt agents such as chatbot agents.
For operators, the useful signal is that model routing, inference economics and deployment control are becoming procurement issues, not only architecture debates. AMD is using that shift to argue for infrastructure that can run a mixed AI estate, although the company did not quantify how its approach changes cost or performance for a specific enterprise workload.
This story draws on original reporting from SiliconANGLE.