• fluxx@mander.xyz
    link
    fedilink
    English
    arrow-up
    16
    ·
    7 hours ago

    It works for them not scraping their own slop back into training data. I assume that is actually the real purpose of the system. They don’t want to share the key with the public. But they probably will with other llm companies in exchange for theirs.

      • 🌞 Alexander Daychilde 🌞@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        ·
        3 hours ago

        The article seemed to claim that some models allow outsiders to submit text for detection if I read it right. That seems like a decent way of doing it. If you “open source” the raw data, it means people can do things to try and get around it - same reason most websites don’t reveal their anti-spam techniques.

    • brsrklf@jlai.lu
      link
      fedilink
      English
      arrow-up
      1
      arrow-down
      1
      ·
      edit-2
      3 hours ago

      That would only work for their own slop though. Anthropic cannot recognize Google’s watermark, only theirs.

      I assumed the goal might be so they could check whether other models have been trained on their output. Like anthropic using that as a “proof” when they start whining again about Chinese “distillation attacks”.

      Of course, since they’re the only ones able to check their watermark, it would be rather shit as evidence anyway. “We’ve run the numbers, and we know you can’t, but trust us, this chatbot is totally copying ours!”