Amazon is reportedly purchasing rare, out-of-print books and physically destroying them to scan their contents for LLM training data to avoid 'model collapse'.
- Amazon is buying rare books, removing spines, and scanning them at a Las Vegas facility.
- The facility, VGT3, is specifically targeted for digitizing non-internet text.
- Rare books provide 'pure' human data, essential for preventing AI model collapse.
In a startling revelation that highlights the aggressive race for data supremacy, Amazon is reportedly acquiring vast quantities of rare books only to dismantle them. According to an investigation by 404 Media, the company is cutting off the spines of these valuable texts to facilitate high-speed scanning, effectively destroying physical literary history to feed its Large Language Models (LLMs).
The operation was uncovered when a tracking device was embedded in a rare book, which eventually led investigators to an Amazon facility in Las Vegas known as VGT3. The site is marked by a distinctive logo featuring a dinosaur clutching a book. In response to the allegations, Amazon stated that it "purchases books through commercial channels to improve the products and services customers use," though it provided no further detail on the destructive nature of the process.
Why This Matters
BozokMedia analysis shows that the industry is hitting a 'data wall.' Most LLMs have already ingested the vast majority of the public internet. To continue evolving, companies like Amazon need high-quality, nuanced text that hasn't been digitized. Rare books, particularly those published before the AI era, represent a goldmine of authentic human thought, free from the contamination of synthetic text.
"The desperation for 'clean' data is driving tech giants to exploit physical archives, treating cultural heritage as mere raw material for computation."
This practice is driven by the fear of 'model collapse.' This phenomenon occurs when AI models are trained on AI-generated content, leading to a degradation in quality and an increase in errors. By sourcing texts from before 2022, Amazon ensures that the training data is 100% human-authored, thereby preserving the intelligence and creativity of the model.
Historically, the digitizing of books (such as the Google Books project) was framed as a way to make knowledge accessible. However, Amazon's approach is fundamentally different; it is not about accessibility, but about proprietary ingestion. The physical destruction of these works means that once they are scanned, the original artifact is gone, leaving the digital ghost in the hands of a corporation.
Frequently Asked Questions
Q1: Why can't Amazon just use digital versions of these books?
A: Many rare books are out of print and have never been digitized, making them the only source of unique, human-written data.
Q2: What is the VGT3 facility?
A: It is an Amazon operation in Las Vegas specifically linked to the acquisition and scanning of physical books for AI purposes.