The Price-Cost Squeeze Closing In on AI
US hyperscalers are pouring $695 billion into AI infrastructure just as Chinese models commoditize the product at 35x lower cost and enterprise customers discover the math doesn't work.
In late May, Nvidia VP Bryan Catanzaro made a quiet observation about his own team's AI usage.
For my team, the cost of compute is far beyond the costs of the employees. — Bryan Catanzaro
The company that sells the chips had run the numbers on itself and found the unit economics inverted: the compute bill was consuming the labor savings. Catanzaro's admission is not an outlier. It is one corner of a three-way price-cost squeeze that has become visible across the industry this summer — and that the industry's own leaders are now confirming out loud. The first force is the cost floor. US hyperscalers are pouring capital into AI infrastructure at a scale without precedent in corporate history. The four largest — Amazon, Microsoft, Meta, and Alphabet — are projected to spend $695 billion in capex this year and $870 billion in 2027 [1]. Alphabet alone raised its annual forecast to $91-93 billion, with 60% of quarterly spend going to servers [2]. Nvidia projects global data center capex will reach $3 trillion to $4 trillion annually by 2030 [3]. The spending is increasingly debt-financed. Morgan Stanley found that hyperscalers doubled their gross leverage ratio from 0.9x to 1.8x in two quarters, and UBS estimates $800 billion in additional US credit has funded AI projects over the past year [4]. Jefferies now identifies the AI sector as the largest issuer of investment-grade debt in the United States [1]. OpenAI is the extreme case. The company has signed $1.4 trillion in computing power deals over the next eight years against anticipated annualized revenue of roughly $20 billion — a 70:1 ratio of committed infrastructure spending to current revenue [5]. Its partners — SoftBank, Oracle, CoreWeave, and others — have borrowed at least $58 billion to fund those data centers, potentially creating $100 billion in debt across OpenAI's backers [5]. The second force is the price ceiling. Chinese open-weight models have matched frontier capabilities and are now setting the market price for inference at a fraction of what US labs charge. DeepSeek's V4-Pro, a 1.6-trillion-parameter model optimized for Huawei Ascend 950 chips, costs 12 to 19 times less than OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 for equivalent tasks [6]. In May, DeepSeek permanently cut V4-Pro prices by 75% [6]. The model now offers outputs up to 35 times cheaper than GPT-5.5 through a combination of permanent discounts and cache-hit pricing [7]. The price gap is not theoretical. In the week ending July 19, Chinese models processed 36.39 trillion inference tokens globally, against 7.39 trillion for leading US models — a roughly five-to-one ratio [1]. Jefferies now describes China's position in unambiguous terms.
China has become a technological peer to the US in AI, as well as in so many other areas. — Jefferies Group
US enterprises are already routing around the price gap. Pinterest uses DeepSeek R-1 for its recommendation engine; Airbnb uses Alibaba's Qwen for customer service. Pinterest's CTO reported that in-house models built on Chinese open-source techniques are 30% more accurate than leading proprietary options [8]. The third force is the customer's breaking point. Enterprise AI adoption is running into a unit-economics wall that the industry's own pricing model created. AI agents — the next frontier of enterprise deployment — consume far more tokens than chatbots, and the bills are arriving. Some enterprise AI token costs have exceeded employee salaries within months [9]. A survey found that 95% of organizations see no measurable return on AI investments, with some experiencing productivity losses from low-quality outputs [4]. The pullback is no longer anecdotal. Microsoft cancelled thousands of internal Claude Code licenses after experimentation costs surged, redirecting engineers to its own GitHub Copilot. Uber exhausted its entire 2026 AI coding tool budget in four months after using leaderboards to incentivize maximum AI usage [10]. Meta CTO Andrew Bosworth reversed internal policy with a direct instruction to staff.
Nobody should be using AI tools just for the sake of using them. — Andrew Bosworth
Meta Chief Scientist Yann LeCun went further, warning that AI labs face a financial reckoning because the costs of running services exceed customer payments.
Labs like OpenAI and Anthropic are going to have to increase prices, they're going to have to cut costs, or there's going to be a big bubble explosion. — Yann LeCun
The phenomenon has acquired a name — "tokenmaxxing" — and it is spreading precisely because the efficiency gains from AI tools, under token-based pricing, paradoxically drive total costs higher [9]. A consultancy called Revenium launched an entire product, AI Outcomes, specifically because enterprises are accumulating what it calls "agent debt" — AI spending incurred without any mechanism to measure the return [11]. These are not three separate problems. They are a single squeeze, each force worsening the others. Chinese commoditization gives enterprise customers a cheaper exit from US closed-model pricing, which deepens the unit-economics rejection inside corporate IT departments, which makes the debt-financed capex — the $695 billion already committed, the $1.4 trillion in OpenAI's deal book — harder to recover. The price ceiling is falling while the cost floor is fixed in concrete and steel. The counter-evidence is real and worth taking seriously. Google Cloud reported a $460 billion backlog with 63% year-over-year revenue growth, including a rumored $200 billion five-year Anthropic commitment [12]. Amundi Investment Institute found that AI stocks lack the "explosive valuation dynamics" of late-stage bubbles, comparing the current cycle to 1995-97 rather than 1999-2000 [13]. Deloitte forecasts that inference will comprise two-thirds of AI computing in 2026, marking a shift from training-heavy spending toward operational deployment [14]. Nvidia is developing a dedicated inference chip to reduce corporate hardware costs, and WEKA and Oracle demonstrated 10x improvements in concurrent users and token throughput [15][16]. But each of these data points confirms rather than breaks the squeeze. Nvidia is building cheaper inference chips because current costs are unsustainable — the company that profits most from AI capex is designing hardware to undercut its own pricing. WEKA CEO Liran Zvibel stated the limitation plainly.
These results prove that AI token economics aren't solved by hardware alone; they're solved by eliminating the memory wall that has been the real ceiling on what existing hardware can do. — Liran Zvibel
Anthropic, despite billions in infrastructure investment, has been secretly reducing Claude session limits during peak hours without formal notice, a rationing move that sparked developer backlash [17]. The Amundi comparison to 1995-97 is instructive in the wrong direction: it means the industry may have years of building still ahead before the reckoning, not that the reckoning has been avoided. And Google Cloud's $460 billion backlog, while real, is concentrated at a single hyperscaler whose own capex is surging — Alphabet raised its forecast to $93 billion — meaning the revenue is being purchased with infrastructure spending that may never earn its way back. The squeeze is now visible to the people inside it. Catanzaro at Nvidia, LeCun and Bosworth at Meta, the analysts at Jefferies — none of them critics, all of them participants — are describing the same geometry from different vantage points. Jefferies, a brokerage whose business depends on capital markets functioning, stated the conclusion in plain terms this week.
The reason credit risk has become more of an issue is that it can no longer be assumed... that large language models will ever be profitable given the related ongoing collapse in token pricing. — Jefferies Group
- 1. Jefferies Warns of AI Capital Destruction Amid Chinese Competition
- 2. Alphabet and Tech Giants Raise AI Infrastructure Spending Forecasts
- 3. Nvidia and Broadcom Lead Massive 2026 AI Infrastructure Boom
- 4. AI Bubble Fears Trigger Tech Sell-Off and Debt Warnings
- 5. OpenAI Leverages Partners for $100 Billion Infrastructure Debt
- 6. DeepSeek Permanently Cuts V4-Pro AI Model Prices by 75%
- 7. DeepSeek Launches V4 AI Model Optimized for Huawei Chips
- 8. U.S. Enterprises Adopt Chinese Open-Source AI Models
- 9. Enterprises Scale Back AI Spending as Token Costs Soar
- 10. Microsoft and Uber Cut AI Tool Use Amid Rising Compute Costs
- 11. Revenium Launches AI Outcomes to Track Agent ROI
- 12. Google Cloud Hits $460B Backlog, Outpaces Rivals in AI Race
- 13. Amundi Report Finds AI Stocks Lack Bubble Dynamics
- 14. AI Computing Shifts Toward Inference and Agentic Orchestration in 2026
- 15. Nvidia Corporation Develops AI Inference Chip to Sustain Growth
- 16. WEKA and Oracle Cloud Benchmark 10x AI Inference Gains
- 17. Anthropic Reduces Claude Session Limits During Peak Hours