OpenAI CFO proposes outcome metric for enterprise AI spend
Sarah Friar’s scorecard shifts AI value measurement toward completed work, as companies face rising pressure to justify model and token costs.
By Dominic Okoye · Staff Writer
· 3 min read
OpenAI CFO Sarah Friar has proposed an outcome-based framework for evaluating AI spending, centered on what she called “useful-intelligence-per-dollar.” The push matters for enterprise buyers and AI vendors because it moves the pricing and ROI debate away from seats and usage volume toward successful work completed at a known cost.
In a Friday blog post, Friar said AI should be assessed by whether it performs meaningful tasks, how much each successful task costs, how reliable the output is and whether the value of each dollar spent improves as usage grows. She also wrote on LinkedIn that traditional software metrics such as seats, active users and renewals are a poor fit for AI, which she said should be measured by work accomplished.
OpenAI did not disclose a new pricing plan, customer adoption data for the scorecard or benchmarks showing how its own products perform under the proposed approach. The company is making a case for a different buying rubric at a time when enterprises are trying to understand whether rising AI budgets are producing operating leverage or just larger cloud and model bills.
The timing is not subtle. Gartner forecasts global AI spending will reach $2.59 trillion in 2026, up 47% from the prior year. That level of spend is drawing more scrutiny from CFOs, CIOs and boards, especially as many projects remain closer to workflow experimentation than measurable financial return.
PwC’s January CEO survey showed the gap. According to PwC, 12% of CEOs said AI had produced both cost savings and revenue gains. Another 33% reported either cost or revenue benefits. A majority, 56%, said they had not yet seen meaningful financial benefits from AI.
That weak payoff picture has made token consumption a board-level concern. Enterprises pay for model usage through processing volume, but the connection between more tokens and better business outcomes can be hard to prove. OpenAI’s proposed scorecard reframes tokens as an input cost that only matters if it produces usable output.
Palantir CEO Alex Karp recently criticized the economics of large language models in a CNBC interview, saying, “The enterprises are just tired of it.” Karp was promoting Palantir’s own products and acknowledged the company has a commercial stake in the argument, so the critique is not neutral. Still, it reflects a broader sales objection facing AI providers: buyers want fewer demos and clearer unit economics.
For OpenAI, the scorecard also serves a strategic purpose. If customers judge AI by completed tasks rather than raw usage, more capable models that handle longer, multi-step work can be positioned as more efficient even when their per-token cost is higher. Friar argued that more advanced models can maintain context, reason through multiple steps, use tools and adapt during a task.
The unresolved issue is measurement. “Successful task” is not a standard accounting category, and different departments will define reliability, quality and value differently. OpenAI has identified the metric it wants buyers to use. It has not shown that the market will accept it as the next standard for enterprise AI procurement.
This story draws on original reporting from CIO Dive.