ThinkPatternGet the app
Perspective
TECHNOLOGY · AUG 26, 2026

AI Is Being Pulled in Two Opposite Directions at Once

The data wall is driving AI labs to buy and destroy physical books, while the power wall pushes models into phones and edge devices — the same industry, moving two ways at once.

A hydraulic blade cuts the spine off a book, and the loose pages feed into a high-speed scanner. This is how the AI industry now acquires its cleanest training data. Amazon's Las Vegas facility and Anthropic's Project Panama are buying millions of pre-2022 books, destroying them, and scanning the text [1][2]. The books are chosen for what they are not: written by humans, before the internet filled up with machine-generated prose. The reason is a problem the industry calls model collapse. Train a model on text another model wrote, and it degrades — irreversibly, in the telling of Serge Gladkoff, who runs the translation firm Logrus Global and warns that freely available human-produced text has been largely exhausted [3]. The free channels are closing at the same time. The UK scrapped its copyright opt-out, Australia rejected a training exemption, nearly 400 newspapers sued OpenAI, and Cloudflare began blocking scrapers [4][5][6][7]. Anthropic paid $1.5 billion to settle a suit over pirated books [8]. Yet two California judges ruled that training on lawfully acquired books is fair use [7]. There are cleaner routes — Amazon is exploring a licensing marketplace, and Google has a method to salvage discarded data — but neither produces new human-authored text [9][10]. So the physical book, bought, destroyed, scanned, became one of the last large sources of clean, human, legally trainable text. Now look in the opposite direction. The same industry is shrinking itself to fit inside a phone. Apple is working to compress 27-billion-parameter models to under 4GB so they run on iPhones [11]. Akamai has built a distributed inference network across 4,400 edge locations [12]. Hugging Face's data shows 83% of model downloads are under 1 billion parameters [13]. Nvidia's newest edge chip is sold on a claim that would have sounded absurd two years ago.

Today’s small and medium frontier models have reached the accuracy of last year’s largest frontier models, unlocking real-time intelligence for edge devices. — Nvidia

The pressure here is the power grid. A 12 GW global capacity deficit is raising the cost of centralized inference [14][15]. The grid is one pressure among several — latency, privacy, and regulation also push inference toward the device [16] — but the energy wall sharpens the economic case for running models where the electricity already is. So one wall centralizes and the other decentralizes, and they are pulling the same industry in opposite spatial directions at the same time. On one side, industrial facilities that disassemble books to extract clean text. On the other, models compressed small enough to live in a pocket. Neither movement should be mistaken for a replacement. The centralized buildout is still the dominant trajectory, and it is accelerating. Nvidia projects $3-4 trillion in annual data center spending by 2030 [17]. Meta forecasts up to $145 billion this year alone [18]. Goldman Sachs sees $7.6 trillion through 2031 [19]. The book scanners and the phone models are hedges running alongside that buildout, not substitutes for it. Which leaves the paradox. A technology sold as immaterial — software, weights, the cloud — is being forced into physical form by the very limits it was supposed to transcend. It is disassembling books to feed itself, and it is shrinking itself to fit inside a phone because the cloud has its own limits: grid, latency, cost. Two directions at once, and neither one resolves the constraint that created it.


Sources
  1. 1. AI Firms Destructively Scan Millions of Books for Training
  2. 2. Amazon Destructively Scans Rare Books for AI Training
  3. 3. Logrus Global CEO Warns of AI Model Collapse
  4. 4. UK Government Scraps AI Copyright Opt-Out Proposal
  5. 5. Australia Rejects Copyright Exemption for AI Training
  6. 6. Nearly 400 Newspapers Sue OpenAI and Microsoft Over Copyright
  7. 7. California Judges Rule AI Training on Books is Fair Use
  8. 8. Judge Approves Record $1.5 Billion Anthropic Copyright Settlement
  9. 9. Amazon Explores AI Content Marketplace for Media Publishers
  10. 10. Google DeepMind Develops Generative Data Refinement for AI Training
  11. 11. Apple Evaluates PrismML Tech for On-Device iPhone AI
  12. 12. Akamai Launches AI Grid Using Nvidia Blackwell Architecture
  13. 13. Hugging Face Data Shows Developers Prefer Small AI Models
  14. 14. AI Data Center Demand Creates 12 GW Global Capacity Deficit
  15. 15. AI Data Center Expansion Hits Power Grid Bottlenecks
  16. 16. Synaptics CPO Advocates for Federated Machine Learning at Edge
  17. 17. Nvidia Projects Trillion-Dollar AI Data Center Spending by 2030
  18. 18. Meta Forecasts Up to $145 Billion AI Data Center Spending
  19. 19. Goldman Sachs Forecasts $7.6 Trillion AI Infrastructure Spend

Keep reading in the app

The full perspective, free in the app.

Download on the App StoreComing soonGoogle Play