ThinkPatternGet the app
Perspective
TECHNOLOGY · AUG 30, 2026

Memory, not compute, is now the ceiling for AI

AI's center of gravity has moved from training models to running them, and the bottleneck moved with it — from compute to memory.

Nvidia shipped three-quarters of its consumer graphics cards this year as the cheap, low-memory versions — 8GB and 12GB models — not because that's what buyers wanted, but because memory chips were scarce. The world's leading compute company was rationing memory, letting memory availability rather than silicon design dictate what it could sell [1].

Current fluctuations in supply for both products are primarily due to memory supply constraints, which have temporarily affected production output and restocking cycles. — ASUSTOR

As a supply-chain story that's just a shortage. It only makes sense once you know the mechanism underneath it. Training a model is compute-bound: the work is matrix math, and the bottleneck is how fast the chips can crunch numbers. Running a model is different. Bernstein's analysis of the inference workload found that its decode stage — the part where the model generates each next word — is memory-bound, not compute-bound. The bottleneck is reading the model's weights and its running context out of memory, not doing arithmetic [2]. So when the industry's center of gravity shifted from training models to deploying them, the binding constraint shifted with it. Lightbits' Abel Gordon put the difference plainly: training is a controlled process, while inference is dynamic and driven by unpredictable real-time demand [3]. A training run is a batch job you schedule; an inference request is a customer asking a question right now, and the model has to hold its entire context in memory to answer. That's why the constraint moved. Once you see that, the same mechanism shows up at every layer of the stack. In silicon, the number more and more chipmakers are competing on is how much memory sits next to the chip. AMD's next flagship, the MI400, is defined by its 432 gigabytes of HBM4 memory rather than by its raw compute [4]. Intel's inference-focused Crescent Island is built around 480 gigabytes of LPDDR5X [5]. Nvidia, the company that built its fortune on training chips, is now developing a dedicated inference chip [6]. Compute hasn't vanished as a constraint — Google's Sundar Pichai says the company is compute constrained in the near term and is turning away customers, paying SpaceX nearly a billion dollars a month for Nvidia GPUs [7]. But that's a company still in the training-era build-out; it's the old constraint, visible next to the new one. In the supply chain, the scramble is more direct. Nvidia committed $279 billion to lock up HBM and server DRAM through 2029 — a figure that doubled in three months — and expects its margins to dip as memory costs rise [8]. Micron sold out its entire 2026 HBM volume and pricing by January [9]. Supply is expanding, not static — Samsung has pushed HBM4 yields to 80 percent [10] — but the fact that everyone is fighting the wall from both sides is itself confirmation the wall is real. The valuations tell the same story. SK Hynix passed Samsung as South Korea's most valuable company this year, the payoff of a 2012 decision to bet on high-bandwidth memory while Samsung stayed in commodity DRAM [11]. On the product side, memory costs are now shaping which products ship: Nvidia raised prices on memory kits to its board partners and put the RTX 50 Super on hold because the 3GB GDDR7 modules were too expensive [12]. And a new software category has emerged specifically to route around the wall. WEKA and Oracle demonstrated a tenfold gain in inference throughput by moving the model's context out of GPU memory into a separate high-performance store [13]. WEKA's CEO Liran Zvibel named the problem directly.

Inference is bottlenecked by how much effective memory is available to GPUs. — Liran Zvibel

The one actor who positioned for this did so a decade early — not by predicting the inference memory wall, but by betting that memory would stop being a commodity and become indispensable infrastructure. SK Hynix chairman Chey Tae-won stated the goal in 2012.

What I really wanted to accomplish when we acquired Hynix was to transform it from a commodity memory producer into a mainstream semiconductor company whose products are indispensable. — Chey Tae-won

That bet is now worth more than Samsung. The $279 billion Nvidia committed — doubled in three months — is what it costs everyone else to arrive at the same discovery: the commodity everyone assumed would scale turned out to be the ceiling.


Sources
  1. 1. Nvidia Prioritizes Low-VRAM GPUs Amid Memory Chip Shortage
  2. 2. Bernstein Identifies AI Memory Bottleneck and Issues Sector Ratings
  3. 3. Lightbits Labs Proposes NVMe SSDs to Scale AI Inference
  4. 4. Broadcom and AMD Gain Ground as Micron Forecasts Memory Shortages
  5. 5. Intel Unveils Three Hardware Architectures for Agentic AI
  6. 6. Nvidia Corporation Develops AI Inference Chip to Sustain Growth
  7. 7. Google Develops Frozen v2 Chip to Solve Gemini Compute Crunch
  8. 8. Nvidia Commits $279 Billion to Secure AI Memory Supply
  9. 9. Micron Sells Out 2026 AI Memory Supply Amid Growth
  10. 10. Samsung Electronics Hits 80 Percent Yield for HBM4 Memory
  11. 11. SK Hynix Surpasses Samsung as South Korea's Most Valuable Company
  12. 12. Nvidia Raises GPU Memory Kit Prices Amid AI Boom
  13. 13. WEKA and Oracle Cloud Benchmark 10x AI Inference Gains

Keep reading in the app

The full perspective, free in the app.

Download on the App StoreComing soonGoogle Play