AI companies are secretly buying, scanning, and destroying millions of physical books to train their models, permanently locking human knowledge inside private corporate servers. Anna’s Archive is urgently calling on volunteers worldwide to scan and upload books to their shadow library before this cultural heritage disappears forever.
The destruction of the book is done to mitigate some copyright concerns. There will be a fairly small number of large-scale projects like this, and each only needs one copy of a given book, so outside of maybe very rare books, the impact is going to be pretty negligible on availability — like someone buying one copy of that book.
Can you explain? The First Sale doctrine means its OK to sell books after you’re done with them. I would think it would be the scanning itself that could be a copyright problem, not the disposition of the physical book after.
I thought the real reason, the main reason, was, as the article says, “Destroying books is cheaper.”
I agree that the concern over the destruction of physical books, and what impact that will have on their availability, is probably exaggerated. Heck, the actual importance of old books to the AI companies may be exaggerated, that seems to be their hype style.
Weird how when I buy a book, the “license” for the content is tied to the physical object, but when a company destroys the book and OCRs it, it’s perfectly ok