Donate the Scan. Except You Can’t.
Used bookstores in the UK are having their best month in years, and the booksellers are not celebrating. The buyers are AI companies, and the books are not coming back.
Here is what is happening. Anything printed before 2022 is guaranteed human-written, which makes old books premium training data. AI labs buy them in bulk, wherever English-language books pile up cheaply, and the UK trade is a rich vein; the books then ship into the legal jurisdiction of the mostly American companies doing the buying. There they are sliced at the spine, run through industrial scanners, and pulped. The destruction is not carelessness. It is the legal foundation. Anthropic learned the cost of shortcuts when a class action over pirated training copies ended in a $1.5 billion settlement, the largest in copyright history, given final court approval this July. The lawful path, blessed by a federal fair-use ruling, requires that one copy exist before the scan and one copy after. A format conversion, not a duplication. The grisly videos of books losing their spines are footage of compliance.
For most books this matters not at all. Shredding one of five million copies of The Da Vinci Code is a loss to no one. But an Edinburgh bookseller told the BBC he has sold the only known surviving example of an 18th-century edition, and books like that are somewhere in the tonnage moving toward the guillotines. When one goes through, the scan survives in a corporate vault. The book survives nowhere.
The knowledge is not destroyed. It is privatized and made uncheckable, which for a scholar is nearly the same thing. A book on a shelf can be found, cited, and disputed. A scan sealed in a training pipeline cannot. If a model tells you what the last copy said, you have no way to know whether it is right, because the only witness was pulped.
The obvious fix seems free: donate the scan. The company already paid for the book and the training use. The file costs nothing to share and, for a rare title, is the difference between existing and not.
Except the law forbids it, and the reason is a genuine trap. The fair-use defense that makes the scanning legal in the first place depends on the scans never leaving the building. Google established this in its decade-long book-scanning case: it won because it showed the public only snippets and kept the full texts locked down. The moment a company hands a complete scan to a library, it is no longer doing protected data analysis. It is distributing copies, and its defense for the original scanning collapses with it. The Internet Archive tested a version of this theory, that owning a print copy entitles you to lend a scan of it, and lost in federal appeals court in 2024. So the good deed is not cheap. It re-opens the exact legal exposure that just cost $1.5 billion to close. Copyright law currently punishes the sharing more reliably than it ever punished the shredding.
The fix that works is three separate fixes, sorted by what the law actually allows.
Audit before you shred. Before a pallet of books reaches the guillotine, check each title, author, and catalog record against WorldCat, the global census of library holdings. If fewer than some threshold of copies sit in public and university libraries, say fifteen, that book gets a non-destructive scan or a donation to an archive instead. This routes around copyright entirely: it is a decision about property, not reproduction, and no publisher on earth can object to a book not being destroyed. It is also the only fix that saves the artifact rather than a picture of it.
Deposit the public domain. Anything published before 1929 carries no copyright at all. There is no legal barrier, none, to uploading those scans to the Internet Archive or the Library of Congress the day after ingestion. No one is doing it. This is the portion of “donate the scan” that survives fully intact, and it is sitting on the table.
Enclave the rest. For in-copyright works, follow the HathiTrust model: deposit scans in a secure academic repository where no one can read or download the books, but vetted scholars can run queries against them. That preserves the closed loop the fair-use defense requires while giving researchers a way to check the model against its sources. It solves the verification problem without touching the distribution problem.
Total cost: a database lookup, some uploads, and an escrow agreement. What it buys: the current story is a video of books having their spines sliced off. The alternative story is a press release about the largest library deposit in history.
I live near chimney stacks from the 1850s, standing alone in the north Georgia woods where the houses around them burned or rotted away. Every one is documented in photographs. The information is safe. Nobody who cares about them thinks the photographs are the point.
Check the shelf before you feed the shredder. It is the cheapest good deed available in AI right now, and the offer expires one guillotine cycle at a time.
Written by human and machines. The legal research in this piece was drafted, deepened, and cross-checked across three AI models, not all of whose contributions made the final cut, and it should still be verified by a person before anyone quotes it in a courtroom.