• 0 Posts
  • 318 Comments
Joined 2 years ago
cake
Cake day: July 10th, 2024

help-circle





  • I do not have the time to work through every part of the example, but imo the main claim is still overstated. Showing that a result would be extremely unlikely under a particular null model is not the same as showing that it is “statistically impossible” for a human to produce. It also does not give a guarantee how the text was written. A tiny p-value is still a probability under assumptions and not a proof of provenance.

    Furthermore, a human does not even need to know the secret key. By pure chance a human written text can display an unusually high alignment with the detector’s secret partitioning.

    The published watermark work, which is also cited by the article, appears to be much more careful about this (based on a quick skim). It reports false positive/negative rates, thresholds, length requirements, and more. Those can be very strong results, provided the assumed conditions apply. They do not turn a detector into an infallible test. Moreover longer text only helps if the assumptions and watermark signal actually remain intact, which can fail in general.

    In such controlled settings, sure, I do not have much issues there. But the claims of “a human cannot replicate this” or “we can guarantee the text was generated with the watermark” are much stronger than the statistics, and especially the cited literature, actually appear to support.


  • Interactive demonstrations are not the same as a formal proof or experimental validation. So we shouldn’t attribute more to this technique than the available evidence can really support.

    I found some time to quickly skim through the sources they have listed. And from that it became pretty clear that this is not realiable in detecting LLM generated versus human output in general. Under very tight assumptions specific error rates were reported that appeared rather low. However, these assumptions do not hold in general, even with more text if no relevant signal remains. There is currently no scientifically validated general purpose way of reliably detection.

    More importantly in the context of Claude, the production watermarking scheme is undisclosed. Therefore, the cited experiments on known watermarking schemes can neither establish how reliably text generated by Claude can be detected, nor how reliably the technique described in the article removes the actual watermark.

    It can be treated as an indicator at best, but not as validated proof.




  • I find the use of this meme template a bit backwards. The one ring is kind of a symbol for evil and corruption. Isildur fell for this corruption and denied Elronds request. A move that can be considered bad for Middle-Earth.

    So placing uBlock in place of the ring, a thing that is wideley considered good from a user’s perspective, is backwards, as this rebellious act of not conforming to the other browser’s changes would then have to be considered bad.

    It would make more sense from Google’s or many other website provider’s perspective. Because for them, ad blockers are really an annoyance. But if you are defending this position, you can be sure to have a mob of angry Lemmy users at your door step next day.

    Thanks for coming to my meme analysis. It was a good waste of time.





  • Although this looks like a clever approach, a kind of stochastic key, I do not see how this guarantees to distinguish text written by big babble machines versus humans. Humans also have a certain pattern of writing, a given distribution of how some words are more likely to appear than others. How can one tell them really apart?

    As an indicator, yeah, might be usable. But I wouldn’t read too much into it before seeing results of a study that runs actual tests.