As AI tech companies increasingly buy and destroy books to feed to their AI models, Anna’s Archive is calling for volunteers to help preserve them for the public record.

  • Truscape@lemmy.blahaj.zone
    link
    fedilink
    English
    arrow-up
    19
    ·
    4 hours ago

    Anna’s Archive does have guides for all of that (apart from the legal advice) on their site (use wikipedia to find the right domain).

    I don’t think you’d want to do so as a university student unless you want to share the same fate as Aaron Swartz though.

    • Catoblepas@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      10
      ·
      4 hours ago

      Surely at least the stuff out of copyright is kosher to scan and share? Which seems like a good area to focus on if they’re being bought up and destroyed

      • dan@upvote.au
        link
        fedilink
        English
        arrow-up
        12
        ·
        edit-2
        3 hours ago

        Stuff that’s out of copyright isn’t an issue though.

        The reason AI companies are destroying books is because current US copyright caselaw (most recently the Anthropic lawsuit) doesn’t allow books to be copied, but considers it fair use if you transform the format of a legally purchased book (like from print to digital) without creating a new copy. By destroying the original book, they haven’t made a copy of it, and so they operate within US law.

        This doesn’t apply to public domain books (books that are no longer copyrighted) as you can do whatever you want with public domain content.

      • Maybe there is a way to do this anonymously? Anna’s Archive deals with books new and old, but maybe that ‘physical-only’ book that has no digital print would be the best place to start?

        I remember no-eyes from #ebooks (IRCHighWay) prioritized recommendations of novels (fiction) to expand their available books, so I think that’s the safest type of book to scan?