ThinkPatternGet the app
Perspective
TECHNOLOGY · JUL 21, 2026

Every AI Giant Is Now Building the Same Chip

Four competitors independently arrived at the same hardware fix for AI's cost crisis, and the physical world cannot deliver it fast enough.

Four companies. No coordination. The same bet. In the past four months, Google, OpenAI, Amazon, and Meta have each independently committed to building their own inference-optimized silicon. Google's eighth-generation TPU splits into two chips: the TPU 8t for training and the TPU 8i for inference, which delivers 80% better performance-per-dollar than the previous generation [1]. OpenAI designed its first custom chip, Jalapeño, in nine months with Broadcom, targeting a 50% reduction in inference cost versus typical GPUs [2]. Amazon's Trainium chip business is now at a $50 billion annual run rate with triple-digit growth, and the company is in talks to sell those chips to third parties [3]. Meta's MTIA program is locked in with TSMC through 2029. None of these companies announced a joint venture or a common standard. They all run the same math. They arrived at the same hardware conclusion. The math is straightforward. Nvidia's GPUs were designed for training: the massively parallel number-crunching that produces a model. Running queries on those same chips — inference — is a different workload with different economics. A training run happens once. Inference happens billions of times a day, and every query costs money. When you are serving AI at the scale these companies intend, the GPU economics that made the breakthrough possible become the economics that make the business impossible. The TSMC and Broadcom pipeline that fabricates these custom chips is now locked through 2031 [4]. That is the detail that tells you this is not a tactical reaction to a bad quarter. It is a structural pivot that will take the better part of a decade to complete. The industry has diagnosed its problem and prescribed the right medicine. The problem is a pricing vise closing from both sides. On one side, enterprises are experiencing what PNC CEO Bill Demchak calls sticker shock as AI providers shift to usage-based token pricing. Uber exhausted its annual AI budget within months. Demchak put the dilemma plainly.

Any impact that AI can have on the productivity of a bank, that productivity can be taken away by the cost of tokens. — Bill Demchak

PNC is now building in-house GPU compute and adopting smaller models to avoid external token costs [5]. It is not alone. DeepSeek permanently cut its V4-Pro API prices by 75% in May, making its models 12 to 19 times cheaper than GPT-5.5 or Claude Opus 4.7 for equivalent tasks [6]. Enterprises facing soaring bills are routing work to the cheaper alternative. That is the demand-side squeeze. On the other side, the hyperscalers are funding $725 billion in 2026 AI capital expenditure through record-breaking bond sales in euros, yen, and sterling [5]. They need higher prices to amortize that debt. But their customers are fleeing to competitors who charge a fraction as much. The custom silicon is meant to break this vise: if you can cut per-query cost by 50% or 80%, you can charge less and still make the math work. The problem is that the silicon takes years to design, fabricate, and deploy at scale, and the repricing is happening now. The market, meanwhile, cannot decide what the right level of AI spending is. In April, Tesla faced investor scrutiny over its $20 billion to $50 billion AI capital expenditure plan. Shares fell 21%. Morgan Stanley warned that investors would need clearer evidence that autonomy was close to support the stock's valuation [7]. By July, the criticism had reversed. Tesla had spent only $2.5 billion of its $25 billion 2026 budget, and investors were pressuring the company to spend more [8]. The same stock. The same year. Two opposite penalties for the same activity. The market is not sending a signal; it is sending noise. This is not because the physical demand for AI compute is imaginary. It is real and it is enormous. Nvidia canceled all 2026 gaming GPU launches for the first time in three decades to prioritize data-center chips, with data-center revenue hitting $51.2 billion of $57 billion in total quarterly revenue [9]. The physical pipeline is so constrained that a 12-gigawatt global data-center capacity deficit exists right now: 8.9 gigawatts operational against 21.1 gigawatts of demand [10]. BlackRock estimates 148 gigawatts of additional power capacity will be needed by 2030 [11]. Jefferies captured the dynamic.

Demand for data centers continues to outpace supply, with hyperscaler capex accelerating and chip volume forecasts implying GWs of capacity ahead of feasible data center delivery. — Jefferies Group

That is the double bind. The industry has correctly identified that the Nvidia training paradigm cannot be the economics that delivers AI at scale. Four competitors independently building the same escape hatch is itself the verdict on the cage. But the escape hatch — custom silicon that slashes per-query cost — requires data centers to house it and power plants to run it. The buildings are not there. The grid is not there. The 12-gigawatt deficit means the chips that would relieve the pricing vise cannot be deployed fast enough to stop the repricing from doing its damage. The diagnosis is right. The prescription is right. The body cannot absorb the treatment in time.


Sources
  1. 1. Alphabet Launches Eighth-Generation TPUs to Rival Nvidia AI Hardware
  2. 2. OpenAI and Broadcom Unveil Jalapeño Custom AI Inference Chip
  3. 3. Amazon Talks to Sell Trainium AI Chips to Third Parties
  4. 4. TSMC Leads Shift Toward Custom AI Silicon and ASICs
  5. 5. Companies Shift to Small AI Models Amid Soaring Token Costs
  6. 6. DeepSeek Permanently Cuts V4-Pro AI Model Prices by 75%
  7. 7. Tesla Faces Investor Scrutiny Over $20 Billion AI Spending
  8. 8. Tesla Faces Investor Pressure Over Low AI Spending
  9. 9. Nvidia Cancels 2026 Gaming GPU Launches Due to AI Demand
  10. 10. AI Data Center Demand Creates 12 GW Global Capacity Deficit
  11. 11. AI Demand Drives Massive Power and Infrastructure Investments

Keep reading in the app

The full perspective, free in the app.

Download on the App StoreComing soonGoogle Play