Tech Brewed

Feeding the Algorithm: Physical Books Sacrificed for the Next Generation of AI

Greg Doig Season 9 Episode 1

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 5:29

Send us Fan Mail

Welcome back to Tech Brewed. I’m your host, Greg Doig, and today we’re diving into a story that feels ripped from the pages of dystopian fiction but is playing out right now in 2026. AI companies are on a buying spree, snatching up physical books by the pallet—especially those published before 2022—just to slice off their spines, scan them, and then shred or pulp the originals, even for rare and irreplaceable editions. Why? Because in a web flooded with AI-generated fluff, these old books are now gold: pure, human, well-edited data for training tomorrow’s machines. The practice is sparking outrage across the board, with critics calling it a modern-day Library of Alexandria disaster. Today, we’ll dig into the trade-offs between digital progress and cultural heritage, the legal green lights making this possible, and the uneasy question: are we preserving knowledge, or erasing its last tangible traces?

Support the show

Subscribe to the weekly tech newsletter at https://gregdoig.com

Hey everyone, welcome back. I'm your host, Greg Doig. Today's episode hit me harder than I expected. We're talking about something that feels like it walks straight out of a dystopian novel, except it's happening right now in 2026. AI companies are buying up physical books, especially older pre-2022 ones by the pallet. They're cutting the spines off, running the pages through high-speed scanners, And then shredding or pulping the originals. We're talking thousands, sometimes hundreds of thousands of books at a time. And the rare ones, the obscure titles with almost no surviving copies left, those are going into the machine too. So let's break it down. The driver here is AI slop. The internet is so full of AI-generated text that companies training the next generation of models are desperate for clean, Purely human-written data. Printed books from before the chatbot boom are gold. They're curated, edited, structured. No machine hallucinations. Enter intermediaries like ISBNDB. They'll source massive bulk orders up to a million books, keep the buyers anonymous with NDAs, and ship them off. Their own marketing even acknowledges the optics problem. One line floating around is basically, AI company destroys 2 million books, is not a headline that generates sympathy. So they coach clients to call it digital preservation. Anthropic is the most documented case. Court filings showed they hired Tom Turvey, the former head of partnerships for Google Books, and set out to get all the books in the world. They bought, destructively scanned, and disposed of millions of print copies to train Claude. A federal judge in the 2025 Bards v. Anthropic case ruled that when you lawfully buy the physical book and destroy it after creating the digital version, that format shifting counts as fair use. Training on those legally acquired copies? Also fair use. Pirated ones? Not so much. So the legal green light is there and the practice is accelerating. So now some reactions. This story blew up fast and the responses have been raw. A lot of people are calling it Library of Alexandria 2.0.

One viral take put it bluntly:

books that survived wars, fires, and centuries of handling are being shredded so an AI can write a better marketing email. Archivists, librarians, and history buffs are furious. Physical copies are artifacts. Once the last few examples of an 18th-century botanical text or some obscure foreign-language economics volume are gone, they're gone for good. You can't reprint the unique physical object. You can't hold the paper that someone hundreds of years ago held. And there's a deeper fear too. If the original is destroyed and only the AI version remains, what stops subtle rewriting or distortion over time? If you destroy history, you can use AI to rewrite it, one commentator said. Another simply posted, Burning of the Library of Alexandria 2.0. Booksellers themselves are conflicted. One told 404 Media that his sales jumped from maybe 20 books a week to hundreds. Great for clearing dusty inventory, especially rare and foreign language titles that weren't moving. But he doesn't like the end use. He doesn't like uncommon books getting pulped. On the other side, some defend it. The knowledge isn't erased, it's digitized and will power better models. Unused books sitting on shelves forever help no one, they argue. Content lives on, just in a different form. Progress versus heritage. That's the tension. I keep coming back to the irreversibility. Scraping the web is one thing. Torrenting libraries is another. This feels different because the physical evidence disappears, and because a judge said it's legal, it's not stopping; it's scaling. So where does that leave us? Should there be carve-outs for truly rare books? Should AI companies be required to non-destructively scan? scan and return or donate the physical copies? Or is this just the cost of feeding the data hunger? So are we preserving knowledge or erasing the last tangible pieces of it? That's the story for today. Thanks for listening. Stay curious, stay skeptical, and I'll catch you in the next one.

Podcasts we love

Check out these other fine podcasts recommended by us, not an algorithm.