The AI Labs Are Selling a Number Only They Can See
The AI labs are pushing a cost-per-task metric that DeepSeek wins by 12 to 19 times when anyone else measures it — and no enterprise customer can check the math.
In May, DeepSeek cut prices on its V4-Pro model by 75%, permanently. When the independent benchmarking firm Artificial Analysis measured what the cut meant in practice — cost per equivalent task, not cost per token — DeepSeek's V4-Pro was 12 to 19 times cheaper than OpenAI's GPT-5.5 or Anthropic's Claude Opus 4.7 for the same work. [1]
It is an efficiency gain being passed through. — Sanchit Vir Gogia
The finding landed awkwardly because the US labs had, by then, begun promoting cost per task as the metric that matters. OpenAI and Anthropic argued that more capable models with expensive tokens could complete work in fewer passes, making them cheaper where it counts. [2] The logic is sound in principle. The problem is that when a neutral party applied it, the labs lost — and lost badly.
This is why the price cut is permanent rather than promotional. — Sanchit Vir Gogia
That gap — 12 to 19 times — is not the whole story. The deeper problem is that no enterprise customer can verify the number the labs give them. The old metric was cost per token. It was crude, but it was checkable: the number of tokens consumed appeared on the invoice, and a customer could multiply. The new metric, cost per task, requires something no enterprise can currently do: trace a single business request through agentic loops to the physical compute consumed. Srikanta Datta, a senior enterprise AI leader at Coupang, gave the problem a name. [3]
If you can't trace a request to the silicon that served it, you aren't operating AI infrastructure. You're guessing. — Srikanta Datta
Without that trace, cost per task is not a number a customer can verify. It is a number the lab reports. A third-party observability firm, Revenium, launched a tool this year explicitly to bridge the gap, and its CEO was blunt. [4]
You are spending now against a return you cannot measure. — Greg Rowell
The labs are promoting a number only they can produce. The customer cannot pick up the ruler. Three pressures are converging in the same window to make that asymmetry matter. First, the IPO calendar. OpenAI and Anthropic both filed confidential IPO paperwork — OpenAI in June, with combined valuations approaching $4 trillion, and OpenAI projecting $600 billion in spending by 2030. [5] More recent filings landed this week. [6] Under cost per token, the labs' operating margins are deeply negative — Apollo's chief economist Torsten Slok calculates that model developers operate at -59% margins while silicon and equipment companies earn 41%. [7] Under cost per task, those margins read as investment, not loss. The metric pivot and the IPO filings serve the same narrative. Second, the customer is squeezing back. Paramount Skydance imposed monthly per-employee spending caps on Claude AI tokens this week — the first visible case of a major enterprise rationing AI access to control costs. [8]
As part of ongoing AI governance, monthly spend limits have been applied to Claude accounts to ensure controlled spending and effective usage across the organization. — Paramount
A metric that promises the bill is actually getting cheaper, even as the invoice rises, lands differently when customers are imposing hard caps. Third, open-weight Chinese models have made per-token pricing obsolete. Developers are releasing high-performance models that users can run locally for free, directly undercutting the proprietary API-access model. [9] The metric shift moves the contest to ground the labs define — a number the open-weight competitor cannot produce because it has no billing relationship with the user at all. There is a contradiction embedded in all of this, and it is visible on the same news feed. Even as OpenAI and Anthropic promote cost per task, they have slashed token prices by 80% on some models. [10] If cost per task is the better frame, why cut the old one? The cuts are a competitive response to Chinese models. The metric pivot and the IPO filings serve the same narrative. The two moves are not inconsistent — they are aimed at different audiences — but they reveal the tension: the labs need to win on price to keep customers, and they need to win on narrative to keep investors. Anthropic has been building a third answer, one that does not depend on trust at all. In February, it launched enterprise plugins that embed Claude directly into Excel, PowerPoint, and Slack, carrying context between applications. [11]
Now, Claude works inside the tools knowledge workers already use — Excel, PowerPoint, Slack — not as a separate window, but as part of how work actually gets done. — Anthropic
The plugins create integration lock-in: once a company's workflows run through Claude inside the tools employees use every day, switching to a cheaper model means retraining staff, rebuilding integrations, and losing the context the model has accumulated across applications. The cost of leaving becomes higher than the cost of staying, even if the customer distrusts the math on the bill. That is the pattern the labs are assembling, piece by piece, in a single season. A metric customers cannot verify. An IPO window that rewards the efficiency story that metric tells. Enterprise plugins that make leaving expensive. The metric shift does not give the customer a better measure of efficiency. It gives the labs a measure only they control — and the customer is being asked to pay what the lab says, measured the way the lab measures, with no way to check the arithmetic.
- 1. DeepSeek Permanently Cuts V4-Pro AI Model Prices by 75%
- 2. OpenAI and Anthropic Push Cost Per Task AI Metric
- 3. Srikanta Datta Identifies Request-To-Silicon Gap in Enterprise AI
- 4. Revenium Launches AI Outcomes to Track Agent ROI
- 5. OpenAI and Anthropic File for IPOs Amid AI Price War
- 6. OpenAI Inc. and Anthropic File Confidential IPO Paperwork
- 7. Credit Markets Signal Doubt Over Trillion-Dollar AI Investment Boom
- 8. Paramount Skydance Limits Claude AI Spending Amid Merger Pressure
- 9. Chinese AI Developers Launch Open-Weight Models to Challenge US Firms
- 10. OpenAI and Anthropic Slash Prices to Counter Chinese AI
- 11. Anthropic Launches Claude Enterprise Plugins and Private Marketplaces