• sachamato@lemmy.world
    link
    fedilink
    English
    arrow-up
    9
    ·
    12 hours ago

    Great that they are digitalising, but can’t they also find a non destructive way to do so? I wonder… Books have always been there and are inherent to our history and human existence. This should not be happening, ever. Books need to be well preserved.

    • dhork@lemmy.world
      link
      fedilink
      English
      arrow-up
      9
      ·
      11 hours ago

      But the destruction is the essential part, ironically, that makes this copying kosher for copyright purposes:

      Destroying books, it turns out, isn’t just cheaper than maintaining them: The presiding judge also ruled that it’s transformative enough to constitute fair use under Section 107 of the Copyright Act.

      So, somehow, destroying the books afterwards makes it all legal. In spite of the fact that no author who wrote a book in all of human history before about 3 years ago had to worry about a chatbot sucking up all of its work, and vomiting it back out, without compensation.

    • Zos_Kia@jlai.lu
      link
      fedilink
      English
      arrow-up
      4
      ·
      12 hours ago

      I think there’s a categorical issue in how this whole story is framed, at least for the examples that have been cited by the booksellers themselves. These books are “rare” because they are low-print and have basically no demand, think obsolete technical books from the 60s or university theses printed at 100 copies. This kind of books are routinely destroyed by booksellers, and even publishers, because keeping inventory is expensive and nobody is buying or reading them. That’s how AI companies are able to buy them for pennies on the dollar, and whatever they don’t buy is probably eligible for recycling or burning by the seller/publisher anyway.

      It’s basically a way to slurp up a few billion tokens for cheap. They’re destroying the copies by chopping off the spine and feeding the loose pages into page-fed scanners, which is way less expensive than those non-destructive scanning robots, but this doesn’t work on really old books from previous centuries as the paper is different and fragile and will jam your scanner constantly.

      What they’re looking for is industrially-produced books with low circulation, not because they are culturally valuable but because they contain prose that hasn’t yet been digitized and that’s what they need for pre-training.

  • lil_baka@ani.social
    link
    fedilink
    English
    arrow-up
    9
    ·
    16 hours ago

    I wonder if there’s anything positive to be said about ISBNDB. On one hand, they’re supplying AI companies, yes, on the other hand, they seem to be digitalizing previously unavailable paper books, which is kinda useful for preservation? Except, I’m not entirely sure it’s possible to properly get this data out of them on consumer scale: they seem to sell API access, but it’s a big question of how usable and affordable it would be to retrieve a digital version of some rare ancient book that could just disappear into oblivion otherwise. Their cheapest subscription is 15$/month, but I have little idea of how what’s listed there translates into volumes of books you can retrieve during this month.

    • Prior_Industry@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      12 hours ago

      Also what if the beancounters decided keeping this digital library is costing us too much so it’s going bye bye.

  • Blaster M@lemmy.world
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    41
    ·
    1 day ago

    very loaded article riding the popularity of ai bad narrative. Doesn’t remotely tell the whole story, only the parts you want to hear to be mad at.