The US Built Its AI Controls for Training. China Won the Inference Race.
Export controls designed to starve Chinese AI of training compute forced China to build custom inference chips that now let its models run at a twelfth the cost of American alternatives.
DeepSeek permanently cut the price of its V4-Pro model by 75% in May. The reduction was not a promotion. It was a structural efficiency gain, enabled by Huawei Ascend 950 processors that replaced Nvidia hardware the company can no longer buy [1]. Those processors exist because US export controls blocked the sale of advanced training GPUs to Chinese firms — and Chinese labs, cut off from the chips they were meant to depend on, built their own. DeepSeek is now developing inference-specific chips explicitly to bypass those controls [2].
chip export controls were a challenge for the company. — Liang Wenfeng
The result is a cost floor American labs cannot match. DeepSeek's V4-Pro now runs 12 to 19 times cheaper than OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 for equivalent tasks [1]. Its V4-Flash model costs roughly three cents per test — more than 100 times less than Anthropic's Claude Fable 5 — though it scores 50 out of 100 on the Artificial Analysis Intelligence Index, well behind US frontier models [3]. Independent benchmarkers still place Chinese models three to nine months behind US frontier systems in raw capability, a gap that has held steady even as the cost gap has widened [4]. The trade-off is explicit: Chinese models sacrifice peak intelligence for a cost advantage that is now structural, not temporary. That trade-off is winning the adoption race. Chinese AI models have outpaced US models in global weekly token consumption for six consecutive weeks, generating 12.96 trillion tokens to American models' 3.03 trillion. All six of the most-used AI models globally are now Chinese [5]. The US-China Economic and Security Review Commission reports that roughly 80 percent of US AI startups now use Chinese open-source base models, with Alibaba's Qwen surpassing Meta's Llama in cumulative downloads [6]. The adoption is not confined to startups. Microsoft is exploring a self-hosted, fine-tuned version of DeepSeek-V4 to power a lower-cost tier of Copilot Cowork, after OpenAI and Anthropic raised prices and shifted to usage-based billing [7]. Pinterest uses DeepSeek R-1 for its recommendation engine and reports that in-house models built on Chinese open-source techniques are 30% more accurate than leading proprietary alternatives. Airbnb runs Alibaba's Qwen for customer service agents [8]. The competition that produced this outcome is not the one the US containment architecture was built to fight. Export controls, Entity List designations, and sanctions threats were designed for the training race — the competition to build the smartest model by restricting access to the advanced GPUs that train frontier systems. But the race that determines who runs the world's AI has shifted from training to inference: who can deploy the cheapest good-enough model at scale. The Commission itself acknowledged that export controls target the "digital loop" of advanced chips for frontier training but are poorly suited to addressing the deployment-driven competition [6]. The intelligence sacrifice Chinese models make is shrinking while their cost advantage is hardening. Stanford's 2026 AI Index found the US-China model performance gap has "effectively closed," with Anthropic holding only a 2.7% advantage as of March [9].
The US-China AI model performance gap has effectively closed. — Stanford University
China now produces 74.2% of global AI patents and leads in research volume and industrial robot installations, while the US still produces more frontier systems — 50 notable models to China's 30 in 2025 [9]. The gap that remains is in raw capability at the frontier. The gap that has vanished is in the utility that drives adoption. DeepSeek's Liang Wenfeng has argued that the inference-first world is eroding Nvidia's CUDA software moat — the proprietary layer that locked developers into Nvidia hardware — because coding agents and tools like TileLang have simplified building AI software across chips [10].
That's where the CUDA moat from Nvidia gets broken because CUDA is no longer a factor in the inference side. — Marshall Choy
The hardware lock that export controls were meant to tighten is the lock the inference shift is breaking. The US response has been self-accelerating in the wrong direction. Anthropic quietly reduced Claude session limits during peak hours, a move analysts read as an effort to push power users from fixed subscriptions toward revenue-guaranteed API consumption — effectively raising costs [11]. The company also launched a Cyber Verification Program that gates Claude access for security research behind professional vetting, while banning China-headquartered companies from its services entirely [12][13]. The strategy is containment through terms of service. Its effect is to make Chinese models — permissive, open-weight, and cheap — the path of least resistance for anyone who needs AI to do work. The US government is no more unified. The Trump administration is considering restrictions on Chinese AI models through procurement rules, Entity List designations, and NSA security advisories [14]. But White House AI adviser David Sacks has called the effort something else.
the leading closed labs, already a duopoly in terms of AI model revenue, want the government to eliminate their open-source competition. — David Sacks
The industry is split along the same fault line. When Moonshot AI released Kimi K3 — the world's largest open-weight model at 2.8 trillion parameters — a coalition of 50 companies including Nvidia, Google, and OpenAI urged the government to avoid restrictions on open-weight models, while Anthropic and Amazon pushed for restrictions on Chinese ones [15].
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window. — Moonshot AI
OpenAI signed the letter protecting open-weight models from government restriction while simultaneously lobbying to block Chinese open-weight models specifically — a position that preserves its own option to open-source while asking the government to eliminate the competition that already does. The US is now scrambling to restrict Chinese models coming in. The House has launched investigations into American firms using low-cost Chinese AI. Treasury Secretary Scott Bessent has threatened sanctions against Chinese model makers [13][16]. But the competitive pressure those models carry was manufactured by the controls already applied going out. The Huawei Ascend 950 processors and the custom inference silicon DeepSeek is now designing — chips that were not supposed to exist — are the reason the models cannot be kept out.
- 1. DeepSeek Permanently Cuts V4-Pro AI Model Prices by 75%
- 2. DeepSeek Develops Custom AI Chips to Bypass US Export Controls
- 3. Chinese AI Firms Spark Global Price War With New Models
- 4. Chinese AI Models Lag Behind US Rivals by Nine Months
- 5. Chinese AI Models Outpace U.S. Rivals in Global Token Usage
- 6. US Commission Warns China's Open-Source AI Threatens US Leadership
- 7. Microsoft Eyes Chinese DeepSeek Model to Cut Copilot Costs
- 8. U.S. Enterprises Adopt Chinese Open-Source AI Models
- 9. Stanford Report Says US-China AI Performance Gap Has Closed
- 10. AI Coding Agents Challenge Nvidia CUDA Software Dominance
- 11. Anthropic Reduces Claude Session Limits During Peak Hours
- 12. TrendAI Deploys Anthropic Claude Model to Automate Vulnerability Research
- 13. US and China Escalate AI Conflict Over Security and Exports
- 14. Trump Administration Considers Restrictions on Chinese AI Models
- 15. Moonshot AI Releases Kimi K3 Open-Weight Model
- 16. U.S. Threatens Chinese AI Sanctions as Hardware Rivalry Intensifies