What happened
An investigation by 404 Media, published on August 17, 2026, has revealed that Amazon purchases vast quantities of books—many of them rare and commercially worthless—and destroys them in the process: bindings are sliced off, pages are mass-scanned, and the volume is shredded. Tracking a hidden AirTag inside a rare copy allowed reporters to follow the shipment to a warehouse in Las Vegas, Nevada, where a team designated "VGT3" carries out this work daily. The logo on the team's door shows a Tyrannosaurus devouring a book.
The company has not mentioned artificial intelligence in its communications. To both 404 Media and Ars Technica, Amazon said only that it "purchases books through commercial channels to help develop and improve the products and services our customers use." The statement neither confirms nor denies the use of books for language model training.
The method: by ISBN, not by title
A crucial detail of the operation concerns acquisition strategy. Booksellers have suspected for more than a year that AI firms were buying entire stocks not by title but by ISBN—the universal code identifying each printed edition. The logic is straightforward: collect the largest possible number of distinct ISBNs to ensure digitization of virtually every book ever published. This explains buying patterns that made no sense to the traditional trade: celebrated editions being overlooked while older books, untranslated, with minimal print runs and little shelf movement, are being purchased in large volume.
At the Las Vegas warehouse, employees told 404 Media they scan the barcode or ISBN before scanning the book. Staff also mentioned that early this year, Amazon nearly ran out of books to scan—so much so that some worried about the warehouse closing. The operation is active and supplying data.
Why old books are so valuable to AI
n The priority for pre-2022 publications has a direct technical cause: anything published since may contain machine-written text, and models trained on model output degrade through a recursive process known as "model collapse." By roughly 2022, virtually all printed books were written by humans. Training a model on text generated by other models causes the quality of outputs to degrade progressively. Pre-internet, pre-chatbot human writing is therefore a finite and increasingly scarce source of high-quality training data.Unsealed court documents revealed that Anthropic codenamed its destructive scanning project "Project Panama," with the stated aim of scanning all books in the world. One of the company's co-founders theorized that training AI on books could teach it "how to write well" rather than mimicking "low-quality internet speak." Anthropic also paid $1.5 billion to settle a lawsuit over pirated copies, but publicly claims it does not buy or destroy rare or antiquarian books.
Amazon, by contrast, has not drawn that line publicly. And it trains its models on materials even closer to its own ecosystem: 404 Media reported the previous week that Twitch will train Amazon's AI on user streams unless users find the setting.
The bookseller reaction
The used and rare book market is divided. Stuart Manley of Barter Books in northern England told the BBC he had never seen anything like the orders in 30 years—one recent bulk purchase matched an entire week's normal sales. He does not lament much: "The world no longer needs five million copies of The Da Vinci Code." There are books that have been unsold online for 20 years that have now moved.
Derek Walker of McNaughtan's in Edinburgh draws a clear line: an academic text printed in 100 copies with 75 already in libraries is no great loss. But "it would be a much more significant problem if an 18th-century edition, surviving to this day, were bought for destruction."
David Tobin of Walden Books in northern London summed up the class's unease: it is pleasant to sell titles that have not moved in years, and it would be sad to see them destroyed. Selling decades-old static inventory for AI training use raises questions that go well beyond commercial comfort.
What is at stake
Destructive scanning is possible because a federal judge, William Alsup, ruled in 2025 that Anthropic's use of purchased books was "exceedingly transformative" and therefore fell under fair use. Buying a book and scanning it is legal. Destroying it afterward is what keeps only one copy in existence—the crucial point of the ruling. In Britain, however, the legal situation is different: the starting point is that both creating the training library and training on it require the copyright holder's permission.
Three questions remain unanswered. Will Amazon explicitly state what its book scanning is for? Will it adopt a rare and antiquarian carve-out, as its competitors claim to observe? And does anyone keep a record of what has gone through the VGT3 scanner—because a digital file in a private training set is not the same as a book that still exists on a library shelf?
Sources: 404 Media, TechCrunch, TNW
✓ Independent sources cross-checked and verified before publishing