AI Firms Destructively Scan Millions of Books for Training
AI developers are purchasing and destroying millions of pre-2022 physical books to secure human-authored training data and avoid AI-generated digital pollution.
Artificial intelligence developers are purchasing millions of physical books, specifically those printed before 2022, to train large language models on human-authored content untainted by AI-generated digital pollution. To accelerate digitization, companies employ destructive scanning, using industrial hydraulic cutters to remove book bindings before scanning pages and pulping the remains. This practice aims to prevent model collapse, a degradation in quality that occurs when AI is trained on synthetic data.
Anthropic PBC drove this trend through Project Panama, an initiative to acquire and destroy millions of volumes for its Claude chatbot. The company used shell identities—including Red Sparrow, Blue Finch, and Green Parrot projects—to mask bulk purchases. While U.S. District Judge William Alsup ruled that the one-for-one transfer from physical to digital copy constitutes transformative fair use, Anthropic recently agreed to a $1.5 billion settlement to resolve separate claims regarding the use of pirated books.
Intermediaries like ISBNdb and Zoom Books facilitated these anonymous acquisitions, with some reports tracing shipments to an Amazon warehouse in Las Vegas. Bookseller Scott Brown estimates that tens of millions of books have been destroyed across the industry, far exceeding the two million linked to Anthropic. While Elon Musk has publicly opposed the practice, instructing his SpaceXAI team to use non-destructive methods, critics and antiquarian dealers warn that the process is permanently erasing rare, foreign-language, and out-of-print historical artifacts.