ThinkPatternGet the app
Perspective
BUSINESS · JUL 31, 2026

AI Poisoned the Internet. Now It's Destroying Books to Fix It.

AI-generated text has contaminated the digital commons, so companies are destroying millions of physical books for clean training data — and a federal judge has ruled the destruction is fair use.

The database that supplies AI companies with books to destroy makes a remarkable promise about its product.

Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. — ISBNdb

The contamination is AI's own output. Over the past three years, AI-generated text has flooded the internet — product reviews, blog posts, comment sections, news-like articles — and models trained on that slurry degrade. Researchers call it model collapse: feed a model its own exhaust and it forgets what human language sounds like. The solution the industry has converged on is to reach past the contamination, back to the pre-AI analog record, and consume it. Physically. The operation is industrial. Anthropic's Project Panama uses hydraulic cutters to shear the bindings off books, high-speed scanners to digitize the pages, and pulp mills to destroy what remains. The company's own description leaves no ambiguity.

Project Panama is our effort to destructively scan all the books in the world. — Anthropic

Meta reached the same destination by a different route. In 2023 the company drew up a plan to spend up to $200 million licensing datasets, then Mark Zuckerberg personally authorized abandoning it for what internal documents call a "fair use legal strategy" — downloading pirated books from Anna's Archive, LibGen, and Sci-Hub instead [1].

We will fight this lawsuit aggressively. — Meta

Google repurposed its Google Books scanning infrastructure — an operation that once promised to preserve the world's knowledge — to feed its Gemini model, and publishers including Hachette, Elsevier, and Scott Turow are now suing, alleging internal documents estimated potential fines of $10 to $100 billion before the company proceeded [2]. Three companies, three paths, one answer: the pre-AI record is there to be taken. Only Elon Musk broke ranks with a different instruction for SpaceXAI.

I've asked the SpaceXAI team to preserve any rare books in a library and scan them the hard way vs just cutting off the spine and scanning — Elon Musk

On July 27, Judge William Alsup gave the practice a legal floor. In Bartz v. Anthropic, he ruled that buying a physical book, scanning it, and destroying the original is transformative fair use. His reasoning turned on a single idea.

every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. — William Alsup

Fair use has always permitted copying under certain conditions — criticism, scholarship, search indexing. Alsup's ruling extends the doctrine from "you may copy" to "you may destroy the original." The digital copy is treated as a functional substitute for the physical book, which licenses eliminating the physical book. The ruling's limits became visible three days earlier, when Anthropic agreed to pay $1.5 billion to settle claims involving 482,000 pirated books [3]. The settlement drew a line the court did not: acquiring books through piracy was unlawful, but training on them — and destroying them — was fair use.

We reached this settlement in 2025, after the court’s landmark ruling that training AI on books is fair use under copyright law — which remains the law today. — Anthropic

The $1.5 billion bought closure on the piracy question while leaving the fair-use shield for training, including destruction, intact. The settlement explicitly does not grant a license for future activities, which means the legal basis for the next round of acquisition remains unsettled — but the principle that destruction is fair use now sits in federal precedent. The rest of the English-speaking world is moving in the opposite direction. In March, the UK government scrapped its proposed AI copyright opt-out after more than 11,000 consultation submissions and public opposition from Elton John and Dua Lipa. The government stated its position plainly.

We agree that copyright should incentivise and protect human creativity. — Government of the United Kingdom

Australia rejected a similar exemption outright [4]. The European Union has launched an antitrust probe. Non-destructive scanning has been demonstrated by Google, Microsoft, and Harvard [5]. The destruction of books is not a technical necessity; it is a cost choice that American law has now licensed, even as the rest of the English-speaking world builds firewalls against the same extraction.


Sources
  1. 1. Publishers and Scott Turow Sue Meta Over AI Training
  2. 2. Publishers Sue Google for Copyright Infringement via Gemini AI
  3. 3. Judge Approves Record $1.5 Billion Anthropic Copyright Settlement
  4. 4. Australia Rejects Copyright Exemption for AI Training
  5. 5. Judge Rules AI's Destructive Book Scanning as Fair Use

Keep reading in the app

The full perspective, free in the app.

Download on the App StoreComing soonGoogle Play