- cross-posted to:
- [email protected]
- cross-posted to:
- [email protected]
The symbol for Amazon’s VGT3, the Las Vegas facility where it scans book for AI training data.
Amazon is buying massive quantities of books, scanning them for AI training data, and destroying them in the process.
A 404 Media investigation was able to reveal Amazon’s book buying operation, which hasn’t been previously reported, by placing a tracking device in a rare book we suspected would be acquired by an AI company for training data, and following it around the country to its final destination.
That final destination was an Amazon warehouse in Las Vegas, Nevada. Amazon employees who work at this location say all they do is receive massive shipments of printed books which they then cut the bindings off in order to scan the books more quickly. The printed book is destroyed in the process. The logo of the Amazon team that works at this warehouse, called VGT3, is a dinosaur, brandishing its teeth and with a book in its hands.
“Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” an Amazon spokesperson told me in a statement.
The world’s AI companies are constantly looking for, and spending extreme resources to locate, more material to train their AI models. With books, that sometimes means destroying them in the process, something that large parts of the public have spoken up against, and which we can now confirm Amazon is doing.
In July, I published a story about booksellers who reported a historical spike in sales starting in the past year. They suspected this spike in sales was due to AI companies acquiring any books they can in search of new training data. Printed books are valuable as training data because a lot of the text they contain is not readily available on the internet, which AI companies have already scraped. The data is also conveniently organized and, if the book was printed before 2022, is guaranteed to be free of AI-generated text, which can make any AI model that is trained on it worse via a recursive process called “model collapse.”
📖 Do you know work at a facility where you scan books? I would love to hear from you. Using a non-work device, you can message me securely on Signal at @emanuel.404. Otherwise, send me an email at [email protected].
These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all. But booksellers couldn’t say for certain who was behind the large purchases because the marketplaces where they sell their books keep the buyers anonymous. When an order comes in, a bookseller ships the sold books to a warehouse operated by the marketplaces, where books are sorted and then sent to the buyer.
In July, one bookseller told me they received a very large order of around 1,000 books on Biblio, one of these marketplaces. The seller agreed to put an Apple AirTag provided by 404 Media in one of the books included in this order so we could see where the book was going. And by extension, which company, AI or otherwise, was behind this massive order. 404 Media granted the bookseller anonymity because they worried sharing this information would harm their business. Biblio did not respond to a request for comment.



While writing my other comment, I thought of a way that might be a little bit closer to a solution for the AI use of copyrighted material. In Germany, afaik we are allowed to make and keep a copy of copyrighted material obtained in any legal way (even borrowing, streaming, etc.) for strictly private and personal noncommercial use. As a compensation, when buying a device that can store or process data (e.g. USB-stick, phone, computer), a small part of the money goes to copyright collectives (“Verwertungsgesellschaften”). Copyright holders can be members of those and get compensation or their work via them (those collectives are also the entities that sell you a license to e.g. play music at public events).
I can imagine a similar system for AI: A part of any subscription, fee, etc. for AI could go to those copyright collectives as a compensation for the material used in training. Compared to obtaining a license once for training (like buying physical books), this has the advantage that the compensation would scale with the actual use of the AI and also the revenue of the AI companies.
I think those systems are far from perfect, but that’s probably a much better approximation to fairness than what we currently have
Would make sense if the AIs made any money. This is kind of the same system we currently have while AI burns millions daily. Dollars, tons of coal, gallons of water reserves, creatives’ livelyhoods… take your pick
At least some money is currently paid for AI and eventually the business has to become profitable. The environmental problem ist of course not approached by this. For the creatives this would give at least some compensation. Problems due to future (creative) work taken over by AI is also not covered by this; that’s “just” what happens when a new technology arises that obsoletes certain work. My suggestion is trying to copensate for creative work somewhat proportionally to its use.
The business has to become profitable… To continue existing. There is absolutely no guarantee of that.
There is, however, overwhelming evidence that unless a major breakthrough is discovered within the next year, the well will dry up. They simply cannot afford to live off of investor hype and vibes forever.