A disturbing investigation reveals that AI companies are acquiring rare books in bulk, scanning them for training data, and then destroying the physical copies. This practice sparks a massive debate over copyright and the preservation of human knowledge.

Loading Video...
  • AI companies are buying rare titles in bulk to convert physical knowledge into digital training data.
  • An Amazon facility in Las Vegas (VGT3) is reportedly cutting bindings and scanning pages.
  • Physical copies are destroyed immediately after the digitization process.
  • Major tech firms like Meta, Google, and OpenAI are facing lawsuits from authors and publishers.

The race for Artificial Intelligence supremacy has entered a predatory phase. A recent investigative report by 404 Media has uncovered a systematic process of what critics are calling 'knowledge theft.' The investigation reveals that rare books are being sourced in massive quantities, not for preservation in libraries, but to be consumed as raw data for AI training models and then discarded.

The red flag was first raised by independent booksellers who noticed a surge in high-volume orders for seemingly random, rare titles. Unlike traditional collectors or academic institutions, these buyers showed a total disregard for pricing, paying premiums without negotiation. To uncover the destination, investigators collaborated with a seller via the Biblio marketplace, planting an Apple AirTag inside a shipment of 1,000 books. The tracker led directly to an Amazon warehouse in Las Vegas, identified as VGT3.

Inside the VGT3 facility, the process is industrial and destructive. Workers reportedly strip the bindings from the books to facilitate high-speed scanning. Once the text is ingested into the digital system, the physical remains of the books are destroyed. Amazon has acknowledged buying books through commercial channels to "improve products and services," though it stopped short of explicitly confirming that the VGT3 facility is dedicated to AI training.

Why This Matters

BozokMedia analysis shows that this practice represents a dangerous shift in how corporate entities view human intellectual history. By destroying the physical artifacts of knowledge, AI companies are effectively monopolizing information. When a rare book is destroyed after being scanned, the public loses a physical record of history, while a private corporation gains a proprietary advantage. This is a transition from 'learning' from books to 'consuming' them.

"The industrial-scale destruction of physical books for digital ingestion is a violation of the social contract between authors and the public archive."

This is not an isolated incident. Anthropic previously faced scrutiny over 'Project Panama,' where millions of books were scanned and discarded. This led to a staggering $1.5 billion settlement with authors. Currently, Meta is embroiled in lawsuits over its Llama models, and Google is fighting claims from publishers like Hachette and Elsevier regarding the training of Gemini.

Company Controversy/Project Current Status
Anthropic Project Panama (Mass Scanning) $1.5 Billion Settlement
Meta Llama Training Data Ongoing Litigation
Google Gemini Training (Publishers) Legal Battle Pending
Did You Know?: The 'Google Books' project of the early 2000s faced similar copyright battles, but the current trend is more aggressive as it involves the physical destruction of the source material.

Frequently Asked Questions

Q1: Is it legal to destroy a book after buying it?
While owning a physical copy allows you to destroy it, the legality of using that content to train a commercial AI model without author permission is the core of current copyright lawsuits.

p>Q2: Why can't AI companies just use digital ebooks?
Many rare and academic books are not available digitally. Physical archives provide a depth of structured, human-verified knowledge that is essential for reducing AI 'hallucinations'.