• Casuls_Die_Thrice@lemmy.zip
    link
    fedilink
    English
    arrow-up
    5
    ·
    edit-2
    12 hours ago

    How tho? Will Claude be looking over our proverbial shoulders now, and if whatever we type just so happens to be word-for-word what it said, it can magically flag it?

    • scrion@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      11 hours ago

      You have a fundamental misunderstanding of how the watermarking works. There is no hidden metadata, but most likely, the statistical properties of the generated text are being changed.

      If you met someone on Halloween with a face mask, and they’d always use the word cromulent in every sentence, you’d probably assume it’s your buddy Mark, who is about the only guy in your circle of friends who does that.

      The model will produce a text where individual words at certain positions, or various n-grams encode a kind of fingerprint that will be an indicator for the text being processed by Claude.

      Like, when people have, like, a specific accent or talk in a certain way, you can totally, like, figure out where they’re from, for sure.

      • AvocadoSandwich@eviltoast.org
        link
        fedilink
        English
        arrow-up
        2
        ·
        11 hours ago

        So isn’t the solution then to just feed the Claude output through a second tiny LLM with the prompt to slightly rewrite the input? I mean you can probably use a 1b model to do that and that can run on practically anything nowadays

        • joe@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          ·
          8 hours ago

          Yes, if the watermarking works like we all assume, significantly editing the output will likely break the watermark. They say this in their press release, as well as that the text must be a certain length to be watermarked. (Probably as a function of how the watermark works.)

          I don’t think this is expected to be 100% foolproof. They certainly don’t make that claim; they explicitly say that the lack of a watermark isn’t conclusive evidence that the text wasn’t generated by their models.