OpenAI and Anthropic Launch New Metric—'Cost-per-Task' to Better Measure AI Costs

Wallstreetcn
2026.08.14 08:53

OpenAI and Anthropic are driving a shift in AI pricing standards from "cost per token" to "cost-per-task," emphasizing that high-premium models can complete tasks efficiently in a single attempt, thereby reducing overall costs. Data from third-party firm Ramp also shows that Anthropic's most expensive model, Fable 5, accounts for only 11% of its Claude software spending, indicating that corporate willingness to pay for high-premium AI has hit a ceiling

As the price war among AI models intensifies, OpenAI and Anthropic are joining forces to rewrite the underlying benchmark for measuring model prices—shifting from "cost per token" to "cost-per-task," aiming to rebuild the value framework for their high-premium models in the eyes of enterprise customers.

In recent weeks, the two US AI leaders have successively voiced their positions. In a July blog post, OpenAI Chief Financial Officer Sarah Friar wrote that while low-cost model tokens are cheap, achieving good results "may require more attempts, more time, or more human review"; Jonathan Pelosi, Head of Financial Services at Anthropic, stated in a recent interview that token cost is merely a "proxy metric" for measuring how much AI computing power is consumed, and that "a better proxy metric for every enterprise is the actual cost of completing a task."

The direct backdrop to this debate over pricing narratives is that vendors such as Meta and SpaceX are squeezing the market with extremely low token prices. According to data from corporate spend management platform Ramp, enterprise adoption of OpenAI and Anthropic has slowed in recent months, with existing customers increasingly turning to cheaper open-source alternatives.

Third-party data sends an even more direct signal: Anthropic's most powerful and expensive widely available model, Fable 5, accounts for only 11% of its Claude software spending. Ramp asserts based on this that the industry has "found a new ceiling for what enterprises are willing to pay for AI."

From "Token" to "Task": Resetting the Pricing Anchor

A token is the unit of data processed by AI models and a way to measure workload. The industry has long assumed that the lower the charge per token, the more cost-effective the software is for customers overall. OpenAI and Anthropic are attempting to overturn this consensus.

Sarah Friar wrote:

"Low-cost models may have cheaper tokens, but achieving good results may require more attempts, more time, or more human review; more capable models have more expensive tokens but can complete the same task in one go."

Pelosi's view echoes this: tokens only measure how much AI is used, not the output. He believes that cost per token is "only useful as a proxy metric," and that enterprises should truly focus on the actual cost of completing tasks.

Low-Price Shock: From "Tokenmaxxing" to Cost Caps

The shift in metrics occurs amid heightened competitive pressure. Many vendors are pushing for lower token costs, attempting to undercut companies like Anthropic that offer stronger models at higher prices. For example, DeepSeek's flagship open-source model released in April has input and output token costs that are only a fraction of those of Anthropic's most advanced products.

Gil Luria, Head of Technology and Research at D.A. Davidson, offered an analogy: "What the US labs want to convey is that using open-source tokens is like using cheap toilet paper—it does get the job done, but you need more of it, and there are other unpleasant side effects."

Companies that once encouraged employees to maximize AI usage through "tokenmaxxing" have imposed limits after facing "bill shocks." Luria stated:

"For many of these companies, internal costs have become prohibitively high."

Most Expensive Model "Performance Not Worth the Price": Third-Party Data Reveals the Ceiling

Although vendors like Anthropic have published their own estimates for cost-per-task, an increasing number of customers are turning to third-party benchmarking firms such as Artificial Analysis and Vals AI for independent assessments. Rayan Krishnan, Co-founder and CEO of Vals, pointed out:

"When labs report their own results, they often use internal benchmarks, making true like-for-like comparisons impossible."

In terms of cost-per-task, models from Anthropic and OpenAI remain the most expensive on the market, with Anthropic's Claude Fable 5 topping the list; however, both also offer more competitively priced low-cost products, such as OpenAI's GPT-5.6 Luna.

What truly dampens the high-premium narrative is Ramp's data: Anthropic's most powerful and expensive widely available model, Fable 5, accounts for only 11% of Claude software spending. "With Fable 5, we have found a new ceiling for what enterprises are willing to pay for AI," Ramp wrote in its report. "Beyond this, higher performance is not worth the price."