• AvocadoSandwich@eviltoast.org
    link
    fedilink
    English
    arrow-up
    2
    ·
    11 hours ago

    So isn’t the solution then to just feed the Claude output through a second tiny LLM with the prompt to slightly rewrite the input? I mean you can probably use a 1b model to do that and that can run on practically anything nowadays

    • joe@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      8 hours ago

      Yes, if the watermarking works like we all assume, significantly editing the output will likely break the watermark. They say this in their press release, as well as that the text must be a certain length to be watermarked. (Probably as a function of how the watermark works.)

      I don’t think this is expected to be 100% foolproof. They certainly don’t make that claim; they explicitly say that the lack of a watermark isn’t conclusive evidence that the text wasn’t generated by their models.