Cisco’s Patel says agentic AI tokenomics is the next budget constraint
Cisco’s Jeetu Patel said persistent AI agents will strain inference budgets, with fewer than 1% of potential users deploying them at scale.
By Dominic Okoye · Staff Writer
· 3 min read
Cisco President and Chief Product Officer Jeetu Patel said agentic AI tokenomics is becoming a primary enterprise budget issue as companies move from occasional chatbot use to agents that run continuously. Speaking with theCUBE’s Dave Vellante during SiliconANGLE’s coverage of AMD Advancing AI 2026, Patel said fewer than 1% of potential users are using agents at scale, leaving substantial room for inference demand to rise.
The comments frame a near-term constraint for AI infrastructure buyers: inference spending may become harder to control as agents generate token usage without waiting for a human prompt. Patel said agents can interact with other agents, creating demand patterns that look different from the request-and-response model enterprises have been testing with chatbots.
Patel argued that if agents deliver a material productivity gain over chatbots, supply shortages could persist for a long period. He did not provide a forecast for token volume, infrastructure spending or Cisco revenue tied to the shift.
What is agentic AI tokenomics?
Agentic AI tokenomics refers to the cost and operational tracking of the tokens consumed by AI agents as they perform tasks, call models and interact with other systems. The point for enterprises is practical: a fleet of always-on agents can turn inference from a usage line item into a recurring budget constraint.
That shift is pushing vendors to sell management, routing and cost-control layers around inference rather than only model access or raw compute. Patel pointed to Cisco Cloud Control, which Cisco describes as a unified management plane for AI infrastructure, as the company’s response to distributed inference across cloud environments, private data centers and endpoint devices.
According to Patel, Cisco Cloud Control is meant to show where inference capacity exists and which agents are consuming tokens. He also said the hybrid model involves AMD handling intelligent routing and compute, while Cisco provides networking bandwidth, security and token-usage visibility. The discussion did not include pricing, customer counts or usage benchmarks for Cisco Cloud Control.
Why smaller AI models are part of the cost argument
Patel said enterprises will need smaller, task-specific models rather than sending every workload to frontier models. Cisco’s example is Antares, an open-weight model family the company released for vulnerability localization in code. Patel said the point is to let companies inspect code without sending proprietary data to the cloud.
He also said Cisco continues to work with Anthropic and OpenAI, so the small-model argument is about workload placement rather than replacing frontier models outright. In Patel’s example, a company might use a smaller model to find 70% of vulnerabilities and use a frontier model for the remaining 30%.
For operators, the implied buying question is whether agent deployments can be routed by task, sensitivity and cost before token consumption gets out of hand. That is a more specific issue than generic AI adoption, and it favors vendors that can see across clouds, data centers and edge devices.
Adoption is still early
Patel said ease of use remains a larger barrier than compute for broad agent adoption. He said the technology is moving quickly but is still far from simple enough for billions of people to activate thousands of agents.
The interview took place during theCUBE’s coverage of AMD Advancing AI 2026. SiliconANGLE disclosed that theCUBE was a paid media partner for the event and said AMD and other sponsors did not have editorial control over the coverage.
This story draws on original reporting from SiliconANGLE.