Skip to main content
Get our mobile app
Download on the App StoreGet it on Google Play

New Investigation Reveals

Amazon Destroying Rare Books for AI Data

Investigation reveals thousands of physical books, including out-of-print and scarce titles, are being purchased, scanned, and discarded to feed text into artificial intelligence models, raising concerns about irreversible cultural loss

חנות ספרים של אמזון

Amazon faces fresh scrutiny over how technology companies gather material to train artificial intelligence systems. According to an investigation published by 404 Media, the company has been purchasing large quantities of physical books—including rare, vintage, and out-of-print titles no longer available in stores—then sending them to a scanning facility where bindings are cut, pages are digitized, and the books themselves are destroyed.

The revelation came after 404 Media embedded an AirTag tracking device in a shipment of approximately 1,000 books. The investigation traced the shipment across several U.S. states before it arrived at an Amazon facility in Las Vegas. Workers at the site described an internal operation identified as VGT3, dedicated to processing physical books: receiving shipments, scanning barcodes, removing bindings, and feeding pages through scanners.

"All we do is scan books," one employee reportedly wrote on an internal Amazon worker forum. According to testimonies cited in the investigation, some workers handle cutting the books apart while others manage intake and barcode scanning.

Amazon responded that it purchases books through commercial channels to develop and improve its products and services. However, the company declined to specify how many books have been acquired, how many were dismantled, how many facilities are involved in the operation, or which AI systems receive the scanned material.

These details raise difficult questions, particularly when the books in question aren't available digitally. For AI companies, a printed book represents a rich data source: complete, edited, diverse texts containing knowledge, research, literature, and historical documentation not always found on the open internet.

Why physical books specifically? In the global race to train language models, data quality has become a strategic resource. Companies seek reliable, diverse texts as free as possible from content generated by other artificial intelligence. Books published before the spread of generative AI tools are viewed in this context as especially valuable—since they were written and edited by humans.

Ready for more?

Beyond that, some older or rare books contain material that doesn't appear online: out-of-print editions, local research, specialized reference works, vintage academic papers, biographies, and various cultural documents. To an algorithm, it's text; to collectors, librarians, and researchers, it's sometimes an irreplaceable item of historical significance.

Booksellers who spoke with 404 Media described receiving unusually large orders that seemed atypical: they didn't focus on a particular genre, sought-after author, or collectible edition, and weren't especially price-sensitive. According to them, the pattern resembled a systematic attempt to acquire titles from across ISBN catalogs—essentially building a massive text database from printed literature.

The story fits into a broader legal and public controversy surrounding the use of books for AI training. Technology companies sometimes argue that legally purchasing a copy and converting it to data for research or development purposes may constitute fair use. Conversely, authors, publishers, and copyright protection organizations contend that training models on their works could harm their livelihoods and control over how their creations are used.

But with rare books, the debate extends beyond copyright. Destroying a physical copy can reduce the number of existing copies worldwide, damage private collections, and erase part of print heritage. A book that isn't rare from an AI model's perspective may be extremely rare from the standpoint of a library, researcher, or reader.

"There are different kinds of value—historical, intellectual, emotional," an anonymous bookseller told the 404 Media investigator. "AI companies care about content as words in sequence. But for others, the book itself is much more than that."

Ready for more?

Join our newsletter to receive updates on new articles and exclusive content.

We respect your privacy and will never share your information.

Enjoyed this article?

Yes (27)
No (1)
Follow Us:

Unmissable content


Loading comments...

Also of Interest