A disturbing investigation reveals that AI companies are acquiring rare books in bulk, scanning them for training data, and then destroying the physical copies. This practice sparks a massive debate over copyright and the preservation of human knowledge.
- AI companies are buying rare titles in bulk to convert physical knowledge into digital training data.
- An Amazon facility in Las Vegas (VGT3) is reportedly cutting bindings and scanning pages.
- Physical copies are destroyed immediately after the digitization process.
- Major tech firms like Meta, Google, and OpenAI are facing lawsuits from authors and publishers.
The race for Artificial Intelligence supremacy has entered a predatory phase. A recent investigative report by 404 Media has uncovered a systematic process of what critics are calling 'knowledge theft.' The investigation reveals that rare books are being sourced in massive quantities, not for preservation in libraries, but to be consumed as raw data for AI training models and then discarded.
The red flag was first raised by independent booksellers who noticed a surge in high-volume orders for seemingly random, rare titles. Unlike traditional collectors or academic institutions, these buyers showed a total disregard for pricing, paying premiums without negotiation. To uncover the destination, investigators collaborated with a seller via the Biblio marketplace, planting an Apple AirTag inside a shipment of 1,000 books. The tracker led directly to an Amazon warehouse in Las Vegas, identified as VGT3.
Inside the VGT3 facility, the process is industrial and destructive. Workers reportedly strip the bindings from the books to facilitate high-speed scanning. Once the text is ingested into the digital system, the physical remains of the books are destroyed. Amazon has acknowledged buying books through commercial channels to "improve products and services," though it stopped short of explicitly confirming that the VGT3 facility is dedicated to AI training.
Why This Matters
BozokMedia analysis shows that this practice represents a dangerous shift in how corporate entities view human intellectual history. By destroying the physical artifacts of knowledge, AI companies are effectively monopolizing information. When a rare book is destroyed after being scanned, the public loses a physical record of history, while a private corporation gains a proprietary advantage. This is a transition from 'learning' from books to 'consuming' them.
"The industrial-scale destruction of physical books for digital ingestion is a violation of the social contract between authors and the public archive."
This is not an isolated incident. Anthropic previously faced scrutiny over 'Project Panama,' where millions of books were scanned and discarded. This led to a staggering $1.5 billion settlement with authors. Currently, Meta is embroiled in lawsuits over its Llama models, and Google is fighting claims from publishers like Hachette and Elsevier regarding the training of Gemini.
| Company | Controversy/Project | Current Status |
|---|---|---|
| Anthropic | Project Panama (Mass Scanning) | $1.5 Billion Settlement |
| Meta | Llama Training Data | Ongoing Litigation |
| Gemini Training (Publishers) | Legal Battle Pending |
Frequently Asked Questions
Q1: Is it legal to destroy a book after buying it?
While owning a physical copy allows you to destroy it, the legality of using that content to train a commercial AI model without author permission is the core of current copyright lawsuits.
Many rare and academic books are not available digitally. Physical archives provide a depth of structured, human-verified knowledge that is essential for reducing AI 'hallucinations'.