The Rent Beneath AI's Price Cuts
The AI labs have learned to automate the cost of their own product, and the profit is moving to the memory, power and silicon no model can make cheaper.
This summer, OpenAI set one of its own models to work on the low-level code that runs on its chips, the routines called kernels. The tuning model, Sol, squeezed roughly 15 percent more efficiency out of OpenAI's own stack, and weeks later the company cut the price of GPT-5.6, its newest model, by 80 percent [1]. That is a company using its product to make its product cheaper, then handing the difference to customers. The markdown went industry-wide all the same. Anthropic brought out a model at half the price of its top system, xAI cut Grok's prices, and Meta released the low-cost Muse Spark 1.2 [2]. Every major American lab spent the summer marking intelligence down. OpenAI has framed its own cut as strategy, not distress.
Our strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost — OpenAI
The company says the goal of the cuts is no longer sheer usage but clear return on every dollar spent [1]. At a Goldman Sachs conference this month, its chief financial officer, Sarah Friar, announced that OpenAI's largest models can now train its smaller ones. Investors treat the capability, which the industry calls recursive self-improvement, as a significant reducer of training costs [3]. The margin math is already public. Anthropic's second-quarter revenue, $11.6 billion, passed OpenAI's $6.7 billion, and Anthropic booked its first operating profit, $559 million. OpenAI's operating loss for the quarter widened to $12.3 billion [4]. OpenAI has also begun selling something besides intelligence. On September 16 it launched Sponsored Agents and a conversational advertising platform, framed as building revenue outside its subscription tiers, already running at a $1 billion annualized rate across more than 40 countries [5]. The bluntest verdict comes from OpenAI's own boardroom. Chairman Bret Taylor has been warning founders off the business his own company is in.
This is the proverbial pickaxes in the gold rush. — Bret Taylor
Taylor calls handmade models fast-depreciating assets and points founders toward applications instead [6]. Every point of efficiency the Sol loop recovers makes a unit of intelligence cheaper, and cheap intelligence does not idle the machines. It fills them. The Lawrence Berkeley National Laboratory, a federal energy lab, projects US data centers will draw 9.5 to 15 percent of the nation's electricity by 2030 [7]. On whether efficiency slows that growth, its finding is blunt.
the scale and growth of computational demand more than offset these efficiency gains, leading to continued increases in absolute electricity consumption. — Lawrence Berkeley National Laboratory
Demand swallows the efficiency. That closes the loop from the other end: the same engineering that deflates the model layer's prices manufactures demand at the layer beneath it. Margin is not evaporating as intelligence gets cheap. It is changing floors, out of what a lab can automate (models, software, the kernels Sol tunes) and into what it cannot: memory, power, custom silicon. Beneath the models sits memory, and the shortage bites hardest in high-bandwidth memory, HBM, the stacked chips that feed an AI accelerator its data. Samsung says the gap between what customers want and what it can make widens rather than closes [8].
Based solely on the demand currently received for 2027, the supply-to-demand gap for 2027 is set to widen even further than in 2026. — Kim Jaejune
The same summer OpenAI marked its newest model down by 80 percent, a company that manufactures physical chips posted margins that make it look like a software firm. Micron's gross margin, the share of revenue left after the cost of making the product, ran 85 to 86 percent, on pricing driven by the memory deficit [9]. The shortage has already reached shop shelves. Sony raised the PlayStation 5's price by as much as $150 [8], and the broader bill for computers, software and accessories is climbing with it.
3%+ per month the rise in prices for computers, software and accessories — The first sustained increases since the early 1980s, as measured by Oxford Economics [10]
The industry's biggest buyers agree on what is binding. Ahead of Nvidia's August 26 earnings, Tim Cook, Andy Jassy and Elon Musk each named rising memory costs and shortages of HBM and NAND flash, the storage-side memory, as the critical constraint on their plans [11]. The shortage now visibly caps the labs' own revenue: on September 10, OpenAI stopped taking new sign-ups for its $200-a-month ChatGPT Pro plan because demand for its Astra model had overwhelmed system capacity [12]. OpenAI's product head, Thibault Sottiaux, described the situation from inside.
We’re pulling all the levers possible to sustain the demand, but I’ve not seen anything like it until now and we went through very steep growth before. — Thibault Sottiaux
The labs, for their part, are behaving as if the shortage never ends. Over seven months last year, OpenAI alone committed to roughly $1.2 trillion of compute under take-or-pay contracts, deals that bill the buyer whether or not it ever takes the capacity [13]. Constellation has locked in 920 megawatts of nuclear power under 18-year contracts and raised its 2026 guidance on data-center demand [14]. Anthropic has committed more than $10 billion to custom silicon, plus 3.5 gigawatts of compute on Google's TPU chips starting in 2027 [15]. The public market is pricing the same shortage as a phase. Constellation raised its guidance and its stock still fell, and Vistra's fell on adjusted earnings up 30 percent [14]. This week Micron took a sell rating while posting record profits, with the analysts behind it pegging the stock's intrinsic value some 17 percent below its market price on the argument that memory pricing normalizes as the deficit eases [9]. That is a serious position, not a reflex. Memory has always cycled: prices for DRAM, the everyday memory chip, rose more than 40 percent at Samsung and 30 percent at SK Hynix, and both companies still missed analyst projections on slower HBM4 shipments, while the industry's shift to long-term pricing agreements may cap how high the peak goes [16]. Both bets settle on the same ground: 2027 into 2028, when the take-or-pay contracts stop being booked backlog and start being actual bills [13]. A sell rating written on Micron's record profits and the labs' signatures are descriptions of the same scarcity. Both are in writing. Only one is right about how long it lasts.
- 1. OpenAI Slashes GPT-5.6 Model Prices to Fight Chinese Rivals
- 2. OpenAI and Anthropic Slash Prices to Counter Chinese AI
- 3. OpenAI Announces Recursive Self-Improvement for AI Models
- 4. Anthropic Overtakes OpenAI in Revenue as Losses Widen
- 5. OpenAI Launches AI-Powered Advertising and Sponsored Agents
- 6. OpenAI Chairman Bret Taylor Warns Against Building Frontier AI Models
- 7. US Data Center Power Demand Projected to Double by 2030
- 8. Samsung Warns Global Memory Shortage Will Worsen Through 2027
- 9. Micron Technology Faces Sell Rating Despite Record AI Profits
- 10. AI Memory Chip Shortages Drive First Computer Price Hikes Since 1980s
- 11. Nvidia Faces Rising Competition Ahead of August 26 Earnings
- 12. OpenAI Pauses $200 Pro Plan Due to Astra Demand
- 13. AI Credit Cycle Risks Compare to 2008 Subprime Crisis
- 14. AI Infrastructure Firms Raise Forecasts Amid Data Center Demand
- 15. Hyperscalers Shift to Custom AI ASICs Over Generic GPUs
- 16. Samsung and SK Hynix Report Slower Memory Chip Price Growth