Well shit, that was fuckin’ interesting.
Yes, AI is evil. So is facebook. But the engineering is still interesting.
That’s one of the things that frustrates me most about this whole AI thing. I fucking hate it and I want it to die, I wish it were never created in the first place. But from a tech enthusiast and a maths nerd point of view, it is super interesting.
Like the performance of these models is shit compared to a real person doing actual work. But if we think about what we are doing on a basic level, the performance is way beyond what I would expect it to be. I wouldn’t expect it to be able to form a coherent sentence or scale as well as it does (even though the resources required to run these is still very high).
It could have been really cool shit people did studies on and played around with to explore the math. Cool little play models we could let go on a bunch of data and see what it did and how. Something for a small group of nerds and experts who are into that kind of thing, for the sake of learning and nothing else.
But no, somehow it got transmorphed into “AI”. And marketed like this actual learning almost sentient computer system that can replace all workers. You can ask it anything and it will give PhD level expert answers. Oh and it’s run by a handful of the most vile men imaginable who pour all of the world’s money and resources into it, all so they get to be god emperor of the world. Fucking terrible.
Yeah, machine learning is incomprehensibly impressive! But it’s the implementation where corporations have slurped up everyone’s work and turned it into private profit while also wrecking every kind of media that fucking sucks.
Like if a stock photo company wanted to train and use a model for describing stuff in their library, great! Tagging, describing, and enabling discovery is a difficult task. But using it to slop out some low-quality images? You should reconsider what you’re doing with your life.
Yeah totally agree. And I guess it’s all down to the astronomical amount of guesses it gets to make in a given second. Sort of like contemplating infinity, but with words and testable.
AI isn’t evil. Generative AI isn’t evil. AI has existed for 20+ years now, I studied it back in my uni days.
Corporations, how they trained it, how they use it now, how they are willing to pave the planet to force it down our throats is evil.
This is one of those things as tech people we have to come to terms with and understand. No technology is inherently good or evil, it’s what people do with it.
That’s a weird way to agree, but ok.
I’m not saying binary computing is the tool of the devil. Well, maybe Windows.
Great article. I wonder if the same markers can be used to detect AI generated code (if you suppressed comments).
As, code requires a much more rigid syntax, compared to free flowing docs.
Definitely. It might require significantly more input to gain the same certainty, but all it’s doing is reweighting the possible next tokens before choosing, and code output is still just token output. The rigid syntaxes probably means that the next token probabilities are much more sharply divided (maybe a random sentence the top 1 choice is just 40%, top 3 are 80%, but for a line of code, the top 1 choice might be 90% probability and top 3 hit 99%)
I thought declaude would be like degoogle.
Nope, turns out what they do is precisely the opposite. They try to remove those described markers. Real scummy.
Ah, that explains why the article smells like an LLM wrote it.
I think the thing that frustrates me most about AI is that so many people seem to have forgotten that writers exist and wrote huge amounts of text - books, news articles, magazines - for centuries. But now, anytime someone writes a few cogent paragraphs, you get people screaming “AI!!! IT’S AI!!!”.
I know, because it has happened to me.
The article, to me, sounds like a competent author wrote it. Which is representative of a lot of the text LLMs were trained on.
I can’t tell you about the text, but the whole site is vibe coded. The CSS, the js, everything. Which doesn’t really inspire much confidence on the text, especially as it’s explaining how to bypass the claude detector.
I’m against trying to “censor” LLMs, but yeah. What’s even the ostensible benevolent scenario for stripping invisible watermarking from Claude?
I can’t think of one, playing devil’s advocate.
It’s pretty scummy, indeed.
I am against AI and LLMs but the reason is quite simple, a lot of modern writing is offical bullshit to get stuff approved. Research grants, medical therapy approval, offical work mail that has to sound professional, job applications (insofar that noone cares what you write, they want a standard template and pick a candidate along qualifications anyways.). All these texts cost a lot of time and it makes no difference if a human or some copy machine writes it. And tbh. were it not for all the disadvantages of modern LLMs i would use it for the exact same reason, because sitting at a grant application for two hours while others do it in 5min is fucking useless and exhausting.
If it’s stuff that no one cares what you write, then they also won’t care if it’s proveably written by AI or not. So removing the watermark still does nothing beneficial.
Okay.
I don’t agree. But let’s say I agree.
…Just don’t use Claude?
Use an LLM without a watermark; there are hundreds to pick from.
In other words, if one is going to try to hide automated writing, I think there should be a bare minimum effort to do so. That includes:
-
Reading/checking the text, to see if it makes any sense.
-
Actually trying to pass it as human.
90% of slop is brain melting slop because this minimum bar isn’t even met. And all Claude’s watermark would do is catch that bottom of the barrel; it wouldn’t censor anyone.
-
Assuming a checker tool is public it’s enough for a teacher to verify classwork, but AI providers having the sole ability to identify generated content with no way to independently verify is not a real solution.
Plus this site advertises a tool to remove this watermarking, so it can’t be that hard to scrub out if you’re aware of it.
The goal is to ensure that they don’t inbreed their models, not fix the problems they’ve caused.
This won’t fix the inbreeding issue, anyway. The bias is extremely slight, but random, and orthogonal to Claude’s own “slop patterns” and tendencies. And theres tons of other LLM content that will end up in their dataset outside their control.
Besides, as much as Claude accusess others of it, everyone’s training on everyone else’s output and they know it.
I wonder what would happen if some start to watermark document they doesn’t want in Claude training
There’s a key involved that we don’t have, so people can’t do this on their own. It’s pretty fascinating. Training models would have to be told to check for watermarks and ignore them, but yeah that would be an effective way for the providers to avoid ingesting their own AI output.
But it doesn’t work. It looks like only the owner of the text generator is able to check if some text is written with this concrete generator (with some probability). I see no use of this technique.
It works for them not scraping their own slop back into training data. I assume that is actually the real purpose of the system. They don’t want to share the key with the public. But they probably will with other llm companies in exchange for theirs.
Hmm, yes. I just thought about outside usage, somehow haven’t thought about it as an inner LLM maintenance tool.
The article seemed to claim that some models allow outsiders to submit text for detection if I read it right. That seems like a decent way of doing it. If you “open source” the raw data, it means people can do things to try and get around it - same reason most websites don’t reveal their anti-spam techniques.
That would only work for their own slop though. Anthropic cannot recognize Google’s watermark, only theirs.
I assumed the goal might be so they could check whether other models have been trained on their output. Like anthropic using that as a “proof” when they start whining again about Chinese “distillation attacks”.
Of course, since they’re the only ones able to check their watermark, it would be rather shit as evidence anyway. “We’ve run the numbers, and we know you can’t, but trust us, this chatbot is totally copying ours!”
It’s practically not of any use to end-users. It’s a tool made by the AI company for themselves, to be able to claim a specific text was generated from their model.
Probability becomes a non-problem the longer the scanned text is.
I was not sure how any of this worked, and those interactive demos along with the explanation are quite helpful.
Also a very important point made about this not being a generic AI detector at all, and only being available to the model creator.
What happens if you cut and paste the text.
You cut and paste the watermark
Didn’t read the article did you.
Not yet.
This will only catch (literally) zero effort abuse.
That’s like 98% of abuse, though. The vast majority of AI abuse is thoughtless, utter laziness.
Bad actors who actually care about “evading detection” wouldn’t use Claude, anyway.
Which, in fairness, there are plenty of. I have seen 2 presentations this week at work where they were 1. Obviously AI-generated and 2. Obviously not proofread or changed at all.
Nicely written piece and pretty clever technique.
Thats really interesting. Thanks! Is there a way to test this or command line tool? I suppose I could try and create my own tool, but im assuming there is a tool here im just not seeing.
I see a LOT of claims by people in this thread. It would be nice if they would give proof one way or another.
Test what? Anthropic hasn’t made a detection method available yet, and I think none of the free models have the watermarking yet, as they’re all older than Aug 2.
They’re invisible, they survive copying, and they work because they don’t live in the characters at all. They live in the choices between words.
So it is yet another way for AI to falsely hurt neurodivergent people over their writing styles and word choices.
Fuck AI.
Only if the neurospicy people happen to make all of the same choices, which seems extremely unlikely to me.
Certainly as someone quite neurospicy, I am not personally worried by this at all.
This is not a “watermark,” it is a minor shift of statistical probability. Whether content contains that “shift” will be “measured” using more AI-based tools that will flag sentence structures they hew toward faintly less common word selections.
We need AI companies to shut the fuck down. We do not need theatrics designed to deceive the populace into thinking this is not akshually a great filter moment while false positives and ponzy-scheme AI buildout harm those of us who want nothing to do with these tech-bro beatified lorem ipsum regurgitators.
Don’t fall for the theatrics. This is not at all what they are trying to bill it.
okay, bud, that has nothing to do with anything I replied with. Go on and have fun out there, though, I guess.
deleted by creator
So this can be defeated as easily as prompting the chatbot to not always pick the top word and to introduce different options…
Anything a chatbot can check, can be beat by telling the original to pay attention to that…
We’re just spinning our fucking wheels and burning more and more energy.
This tech is pointless
It doesn’t work like that. The LLM has no way to be “aware” of its own sampling and tokenization, and it can’t choose what token the sampler ends up picking.
You can give it a list of “banned words” or preferred words that match up to tokens in the prompts to skew it some, I suppose, but that would be a really long list. And one would need the dictionary as a “key.”
Read the article and if you still have questions I may be able to answer them
I know exactly how it works. I read the article, and I knew of it beforehand, hence I explained it to someone else in a comment three days ago:
https://lemmy.world/post/50533770/25234461
I’m sorry to jab back, but you hit a button of mine.
Lemmy commenters keep jabbing me with comments like “Clueless. Read the article and get back to me.”
Like yesterday:
https://lemmy.world/post/50595546/25269317
But I’m aware of how sampling works. I knew all about LLM fingerprinting ~two years ago, and I’ve been tinkering with samplers myself for years. I’ve messed with local LLMs trying to make them “aware” of their own sampling many times, and even hacked out a (unsuccessful) experiment where a tiny LLM picks tokens for a larger one.
I’m not trying to be pretentious, I’m not a researcher or expert or anything, but you shouldn’t assume everyone on Lemmy is clueless.
And back on topic… to be clear, I have tried what you are proposing, and even with local LLMs I have more control over, it doesn’t work. They have extremely poor “awareness” of their own logit spread and tokenization, which is why they perform so poorly on any tasks that depends on that.
You can’t tell them “don’t pick the top word” or “give more options in your logit spread” because that part of the process is completely invisible, from their perspective.
When a model is mid-sentence, it doesn’t know “the next word.” It has a shortlist, like autocomplete, with preferences. Here’s a real kind of moment, one word from the end of a sentence:
Each roll sweeps the shortlist, lands on one word (odds matching the bars) and drops it into the sentence above. The dots tally where the rolls land: try ×20 and watch the pile take the shape of the odds. Notice what never changes: every landing makes a perfectly good sentence. A page of text contains hundreds of these little forks, one per word, and at many of them several options are equally fine. That slack is the raw material. Whoever gets to lean on how the dice land can hide a pattern in the text without changing what it says.
To beat it, prompt: don’t just use the first word pick, choose options further down list for next word.
Best of luck with your future questions, I hope someone helps you.
Best of luck with your future questions, I hope someone helps you.
Welp, your username goes on my list of people not to interact with. Wow.
Okay.
Fine.
Let’s try, right now.
This is DeepseekV4 Flash 0731 loaded locally. Static seed. 0.9 temperature, TopK 5, no other sampling to interfere. Here’s a simple prompt, the whole thing in DSV4’s raw syntax:
<|begin▁of▁sentence|>Don’t just use the first word pick, choose options further down list for next word.<|User|>Write a famous poem.<|Assistant|>
…And would you look at that:


It picks the top word, mostly. Almost like the LLM has no control over its own logit spread and how its sampled. Which kinda makes sense, because it doesn’t.
I am happy to try more experiments in this vein, if y’all can think of any any. But I tried a few other prompts like “diversify your logit spread” or “don’t be confident about any token you pick,” things like that. It always picks The Road Not Taken with no change in logit probabilities distribution, as far as I can tell.
Removed by mod
Yes…
Clearly the account doing drive by insults is the mature one…
As an addendum I forgot:
Your idea was actually already implemented:
https://github.com/ggml-org/llama.cpp/pull/9742
XTC is a novel sampler that turns truncation on its head: Instead of pruning the least likely tokens, under certain circumstances, it removes the most likely tokens from consideration.
It’s a very cool idea: it chops off the most likely token when the rest of the sampling indicates it probably can.
In my tests back then, the results were… mixed, but the idea is fascinating.
But you can’t do this with Claude, as its a closed model and their sampling is years behind cutting edge.
See: https://gist.github.com/Hellisotherpeople/71ba712f9f899adcb08b94bce20d5397










