Booksellers had been noticing a strange pattern for about a year: someone was regularly buying up rare and low-print-run books in bulk, with little regard for their content. Suspicion fell on companies training large language models — but there was no proof. Now, it seems, proof has been found in a highly unusual way.
A tracker in a book led to an Amazon warehouse
One seller decided to test the theory and slipped a hidden AirTag Bluetooth tracker into a book from a wholesale batch. The device revealed the route: the order went not to a private collector and not to an online store's warehouse, but to Amazon's AI training center in Las Vegas. A team there works by literally tearing the spines off books and scanning the pages — a method that lets them process volumes faster and more carefully than a regular scanner.
An investigation published by 404 Media confirmed that the bulk purchases of rare books are indeed linked to Amazon. The warehouse with the code VGT3 houses a digitization team, and the facility's door is decorated with a logo of a tyrannosaurus about to eat a book. Fitting, given the fate of the copies that end up there: after scanning, the books are destroyed.
Amazon itself declined to comment on the findings, limiting itself to a formal statement: the company buys books through ordinary commercial channels to "develop and improve products and services for customers." AI training is not mentioned in that response, even though that appears to be exactly what the hunt for unique texts is for.

Why Amazon destroys rare books
Amazon is actively building its own frontier models — a new generation of AI that is meant to compete with the work of Google, OpenAI, and Anthropic. Such models need enormous volumes of unique data, and the less "worn-out" the material, the better. Ordinary internet content is no longer enough for this purpose.
That is why the company's interest has focused not on expensive first editions, but on books that look unremarkable from a market standpoint: works never translated from rare languages, long out of print, or simply not popular enough for a mass print run. They can be bought in bulk without regret, and they provide what the algorithms need — text that has never been digitized before.
Warehouse employees described their work matter-of-factly on forums: monotonous, but with convenient hours, and adjustable to one's needs if necessary. They were trained to scan the barcode or ISBN before processing each book. This detail is especially important: it indicates that Amazon is methodically working through the catalog of printed publications using a list of identifiers, trying to cover as many works as possible.
According to 404 Media, at some point this year the team ran out of books to scan. The shortage was so acute that workers feared the warehouse would close. Later, deliveries resumed, and the facility continues to operate.

The question of value and ethics
The destruction of books for AI training sparked a strong reaction online. Michael Barry, a well-known journalist and writer, called the practice "the embodiment of evil." Reddit users debated how justified such treatment of cultural heritage is. One argument sounds particularly sharp: many of these books had been gathering dust on store shelves for decades, and hardly anyone would have bought them for reading. But that does not negate their value.
Old editions have more than just monetary worth. There is historical and intellectual value, and sometimes sentimental value — for bibliophiles or researchers. As the bookseller who ran the tracker experiment noted, for AI companies such a book is merely "content in the form of a set of words." Small details of daily life, outdated terms, peculiarities of regional languages — all of this disappears forever after the original is scanned and destroyed.
Skeptics called the publication "propaganda from AI haters," but even they agreed that regretting the loss of unique printed sources is entirely reasonable. Especially given that the contents of the scanned books are unlikely to be made public by the company. The fewer competitors who can access unique data, the better for those who collected it. Libraries and researchers will most likely not get access to such archives.
Unlike Amazon, competitors Anthropic and xAI have already publicly stated that they do not use rare or antique books for training. Amazon, meanwhile, prefers to stay silent, which only raises more questions. If the practice is truly systematic, the consequences for the book market and cultural heritage could be serious: rare editions are a finite resource, and the list of ISBNs that can be methodically checked off is right there in front of the AI giants.



