• sachamato@lemmy.world
    link
    fedilink
    English
    arrow-up
    9
    ·
    13 hours ago

    Great that they are digitalising, but can’t they also find a non destructive way to do so? I wonder… Books have always been there and are inherent to our history and human existence. This should not be happening, ever. Books need to be well preserved.

    • dhork@lemmy.world
      link
      fedilink
      English
      arrow-up
      9
      ·
      12 hours ago

      But the destruction is the essential part, ironically, that makes this copying kosher for copyright purposes:

      Destroying books, it turns out, isn’t just cheaper than maintaining them: The presiding judge also ruled that it’s transformative enough to constitute fair use under Section 107 of the Copyright Act.

      So, somehow, destroying the books afterwards makes it all legal. In spite of the fact that no author who wrote a book in all of human history before about 3 years ago had to worry about a chatbot sucking up all of its work, and vomiting it back out, without compensation.

    • Zos_Kia@jlai.lu
      link
      fedilink
      English
      arrow-up
      4
      ·
      13 hours ago

      I think there’s a categorical issue in how this whole story is framed, at least for the examples that have been cited by the booksellers themselves. These books are “rare” because they are low-print and have basically no demand, think obsolete technical books from the 60s or university theses printed at 100 copies. This kind of books are routinely destroyed by booksellers, and even publishers, because keeping inventory is expensive and nobody is buying or reading them. That’s how AI companies are able to buy them for pennies on the dollar, and whatever they don’t buy is probably eligible for recycling or burning by the seller/publisher anyway.

      It’s basically a way to slurp up a few billion tokens for cheap. They’re destroying the copies by chopping off the spine and feeding the loose pages into page-fed scanners, which is way less expensive than those non-destructive scanning robots, but this doesn’t work on really old books from previous centuries as the paper is different and fragile and will jam your scanner constantly.

      What they’re looking for is industrially-produced books with low circulation, not because they are culturally valuable but because they contain prose that hasn’t yet been digitized and that’s what they need for pre-training.