Jul 31, 2026
Enterprise

AI cost per token loses value as production workloads vary, CoreWeave exec says

CoreWeave executive Chen Goldberg says token pricing misses GPU utilization, latency and elasticity costs in production AI.

Dominic Okoye

By Dominic Okoye · Staff Writer

· 3 min read

AI cost per token is becoming a weaker way to judge infrastructure spending as enterprise systems move from experiments into production, CoreWeave executive Chen Goldberg argued in a guest post for CIO Dive published July 31. Goldberg, CoreWeave’s EVP of product and engineering, said identical GPUs, models and configurations can produce different economics when workloads differ, making a single token price a thin basis for vendor or architecture decisions.

The argument is not that token pricing is useless. Goldberg said it remains easy to compare and discuss, which explains why it became a common shorthand during the first wave of AI adoption. His point is narrower and more useful for operators: published token prices tend to reflect controlled runs with stable workloads and optimized settings, while production systems face latency targets, bursty demand and bottlenecks that change with the application.

Goldberg did not disclose new CoreWeave pricing, customer spending levels or performance benchmarks. The post is best read as a vendor-side view of how AI infrastructure should be evaluated as spending scrutiny rises and companies try to turn AI pilots into systems that handle real workloads.

Why is cost per token not enough for AI pricing?

Goldberg said the same infrastructure can become compute-bound or memory-bound depending on workload shape. Short prompts served in large batches can keep GPUs busy with computation, while long-context jobs may be limited by memory bandwidth because the system has to move and manage key-value cache data.

Latency requirements add another constraint. According to Goldberg, an inference setup that posts high throughput in a benchmark may deliver much less usable output when an application needs real-time responses. That distinction matters for founders and operators selling AI products, because the customer experience is tied to response time as much as raw model output.

The variables Goldberg cited include GPU selection, quantization, attention kernels, cache layout, batching strategy and speculative decoding. Each can affect speed, cost and sometimes model quality, and the right setting depends on the workload rather than a universal best practice.

What CoreWeave says buyers should measure instead

Goldberg argued that companies should focus on how much paid GPU time produces useful work. In his framing, waste can come from cluster stragglers, slow detection of degraded infrastructure and demand patterns that force teams to pay for idle peak capacity.

He described time-to-detection as an economic variable because production systems degrade somewhere at scale, even when vendor price sheets do not show a line item for the operational cost. He also said inference traffic, agents and fine-tuning jobs can arrive unevenly, making elasticity part of the real bill.

For infrastructure buyers, the practical recommendation was to test their own workloads under production-like conditions rather than rely on benchmark tables or chip specifications alone. Goldberg said teams should observe cost, latency, accuracy, component failures and queue behavior, and use that evaluation to decide which model and infrastructure setup fit the job.

The post also signals how AI cloud vendors want sales conversations to change. Instead of presenting compute as a menu of fixed SKUs, Goldberg said stronger vendor relationships now resemble joint engineering work before a deal is signed. That claim serves CoreWeave’s position in the market, but the operational issue is real: token prices are legible to finance teams, while production AI costs are often determined by utilization, failure modes and traffic shape.

This story draws on original reporting from CIO Dive.

More from Enterprise

All Enterprise →