The AI Race Has Split in Two — Over Who Pays to Run the Models
Inference costs are now the industry's binding constraint, and they're pushing US labs to own the whole serving stack while Chinese labs push serving onto users' own hardware.
OpenAI's inference expenses quadrupled in 2025, and its adjusted gross margin slipped from 40 percent to 33 percent [1]. CFO Sarah Friar has said the company worries about whether it can honor future contracts if revenue doesn't keep pace with its infrastructure commitments [2]. That margin compression is the financial pressure behind the wave of moves that followed: the custom chips, the acquisitions, the throttling, the price cuts, and the open-weight countermove now coming out of China. The bottleneck has moved from training to serving. QumulusAI's chief executive put the distinction plainly.
Inference workloads have very different performance and economic requirements than model training environments. — Mike Maniscalco
[3] The American response has been to pull the entire serving stack in-house, layer by layer, all at once. OpenAI and Broadcom built the Jalapeño inference chip in nine months and claim it cuts inference costs roughly in half [4]. Anthropic is standing up its own silicon team [5] and talking to Samsung about manufacturing custom 2-nanometer inference processors [6]. Amazon is steering its $220 billion buildout toward its own Trainium chips, which it says save 20 to 30 percent on inference [7]. Anthropic is also negotiating a $6 billion acquisition of Decart, a startup that squeezes efficiency out of chips across Nvidia, Google, Amazon, and AMD hardware [8]. It has quietly cut Claude session limits at peak hours, a compute-cost move that analysts say may double as a nudge to shift heavy users onto API billing that guarantees revenue [9]. And OpenAI cut Luna's price 80 percent. None of these is a model breakthrough. Each is a layer of the same stack being pulled in-house to make serving cheaper. China is answering the same pressure from the opposite direction. Rather than own the serving stack, Chinese labs are giving the models away and letting users run them on their own hardware. Alibaba's Qwen open-weight models hit 3 billion downloads in six months, seven times Alphabet's 418 million [10]. Hugging Face's verdict on what that means is blunt.
Qwen has become part of the default workflow for developers deciding what models to fine-tune and deploy. — Hugging Face
The open-weight push relocates inference off the lab's servers and onto the user's machine [11]. This is not a clean story of China abandoning training. ByteDance is training a 10-trillion-parameter model [12], and Zhipu just opened a 1-gigawatt data center running on domestic silicon [13]. The divide isn't about who trains. It's about who pays to serve, and on that question the two sides have made opposite bets. The stakes are the depreciation schedules behind the centralized buildout. Amazon's $220 billion forecast assumes inference keeps running in its data centers [7]. But open-weight models paired with local hardware like Nvidia's DGX Spark point toward inference migrating to the edge, and if centralized demand grows slower than projected, that infrastructure may not earn back its capital on the timeline the buildout assumes [14]. The two strategies aren't just different. Each is a bet that the other's economics don't work.
- 1. OpenAI Cuts Infrastructure Spend to $600 Billion Ahead of IPO
- 2. OpenAI Raises Computing Infrastructure Spending Projection to $750 Billion
- 3. QumulusAI Secures $124 Million in GPU-as-a-Service Agreements
- 4. OpenAI and Broadcom Unveil Jalapeño Custom AI Inference Chip
- 5. Anthropic Builds In-House Team to Develop Custom AI Chips
- 6. Anthropic Discusses Custom 2nm AI Chips With Samsung Electronics
- 7. Amazon Raises 2026 AI Spending Forecast to $220 Billion
- 8. Anthropic Negotiates $6 Billion Acquisition of Decart AI
- 9. Anthropic Reduces Claude Session Limits During Peak Hours
- 10. Alibaba Qwen AI Models Reach 3 Billion Global Downloads
- 11. Chinese AI Developers Launch Open-Weight Models to Challenge US Firms
- 12. ByteDance Trains 10 Trillion-Parameter AI Model to Rival Anthropic
- 13. Zhipu AI Launches 1-Gigawatt Data Center Using Domestic Silicon
- 14. AI Data Center Expansion Hits Power Grid Bottlenecks