The race to build the world’s most powerful artificial intelligence models is no longer happening only inside data centers. It is also unfolding inside warehouses stacked with second-hand books.
Court documents, investigative reports, and recent copyright lawsuits have revealed that major AI companies, including Anthropic and allegedly other leading developers, have spent millions of dollars purchasing physical books in bulk, slicing off their bindings, scanning every page into digital archives, and destroying the original copies. The practice, known as destructive scanning, has ignited a global debate about copyright, cultural preservation, and whether humanity’s written knowledge is becoming raw material for AI.
What was once a quiet digitization process has now become one of the most controversial chapters in the AI industry’s rapid expansion.
The Mission: “All the Books in the World”
One of the most revealing disclosures emerged from the copyright lawsuit against AI startup Anthropic.
In early 2024, Anthropic launched an internal initiative codenamed Project Panama, a confidential effort to build a vast digital library for training its Claude AI models. According to unsealed court records, the project’s goal was to “destructively scan all the books in the world.” To achieve this, the company hired former Google Books executive Tom Turvey, purchased millions of second-hand books, removed their bindings with industrial cutters, scanned every page, and recycled the physical copies. Internal planning documents also instructed employees not to discuss the project publicly, highlighting the secrecy surrounding the operation.

Instead of negotiating licensing agreements with thousands of publishers, Anthropic reportedly turned to the second-hand book market.
Millions of used books were purchased from online booksellers and bulk distributors. Once delivered to scanning facilities, workers removed the bindings using industrial hydraulic cutters so every page could pass through high-speed scanners. After digitization, the paper was recycled instead of being preserved.
For AI developers, the objective was straightforward: transform physical books into machine-readable text capable of training increasingly sophisticated large language models.
Why Physical Books Matter More Than the Internet
At first glance, the internet might appear to provide limitless text for AI training.
However, developers increasingly see books as far more valuable.
Books generally contain professionally edited writing, richer vocabulary, longer narrative structures, historical context, and fewer factual errors than online content. They also predate the recent explosion of AI-generated material flooding websites, blogs, and social media.

Researchers and industry observers have noted that older printed books offer “clean” human-authored language that has not been contaminated by machine-generated text, making them especially valuable for improving writing quality, reasoning, and long-form responses.
As generative AI produces ever more online content, obtaining reliable human-written material has become a strategic priority.
Why Destroy the Books?
The destruction of books has shocked many readers, but the motivation is largely practical.
Non-destructive scanning methods, similar to those used by libraries and museums, require careful page-turning and specialized equipment, making them significantly slower and more expensive.
Destructive scanning removes the spine entirely, allowing loose pages to pass rapidly through automated scanners. This dramatically increases throughput while reducing labor costs.
For companies digitizing millions of volumes, speed becomes a competitive advantage. In effect, physical books become temporary containers for data.
The Copyright Battle
The book scanning operation surfaced during one of the most significant copyright cases in AI history.
Authors accused Anthropic of using copyrighted books without permission to train Claude, its flagship AI assistant. The lawsuit also alleged that the company maintained a massive repository of pirated digital books alongside legally acquired physical copies.
In an important legal distinction, U.S. District Judge William Alsup ruled that scanning legally purchased physical books for AI training could qualify as fair use because the process transformed the books into internal training data. However, the court separately found that maintaining pirated copies was not protected under fair use, leading to further litigation and eventually a multibillion-dollar settlement over those unauthorized copies.
The ruling has become one of the most influential legal precedents shaping AI copyright law in the United States.

Not Just Anthropic
Although Anthropic’s internal documents offered an unusually detailed glimpse into the process, court filings and investigations indicate that the broader AI industry has aggressively pursued large collections of books.
Reports involving Meta, Google, OpenAI, and other companies describe extensive efforts to acquire massive text datasets, sometimes through licensed collections and sometimes through legally contested sources. Internal communications disclosed in various lawsuits suggest companies believed obtaining publisher permission at scale would be impractical, prompting alternative acquisition strategies.
Recent investigations also suggest that specialized intermediaries have quietly purchased books on behalf of AI companies to avoid attracting public attention, while online booksellers have reported unusual demand for obscure and out-of-print titles.
Cultural Concerns Beyond Copyright
The controversy extends well beyond intellectual property.
Books are not merely collections of words. Many physical copies contain handwritten notes, historical bindings, unique print editions, or limited publications that cannot easily be replaced.
Critics argue that large-scale destructive scanning risks permanently removing culturally valuable editions from circulation.
Libraries and archivists have long digitized books while preserving originals. The AI industry’s industrial-scale approach raises a different question: should rare physical books become disposable once their text has been extracted?
For preservation experts, the answer is far from settled.
Publishers Push Back
The publishing industry has become increasingly vocal.
Authors argue that AI companies are benefiting commercially from decades of creative labor without permission or compensation.
Publishers have likewise filed lawsuits claiming that books supplied for traditional services such as digital previews or searchable archives were never intended to become training material for commercial AI systems.
Meanwhile, new academic research suggests AI-generated books are rapidly expanding across self-publishing platforms, adding further pressure to an industry already grappling with questions of originality, authorship, and market value.
A New Kind of Resource Race
The scramble for books reflects a larger reality in artificial intelligence.
Just as previous technological revolutions competed for oil, minerals, or computing power, today’s AI race increasingly revolves around access to high-quality data.
Books once valued primarily as cultural objects are now treated as premium datasets capable of teaching machines how humans write, reason, explain, and tell stories.
Whether future AI models continue relying on destructive scanning, shift toward licensed publishing partnerships, or face tighter regulation remains uncertain. What is clear is that books have become one of the most valuable commodities in the AI era, not for their shelves, but for the knowledge embedded within their pages.