As I’ve said in response to a number of articles about this, all this means is that when someone wants to use Claude’s output for something, they just need to remember to manually type it in now, rather than just copy-paste it. Practically speaking, this changes nothing.
What I’m saying is that if someone just types verbatim what Claude says, instead of just going “Ctrl-C, Ctrl-V”, that completely undermines this whole thing.
Their point is that Claude is now engineering its responses to specifically conform to a detectable pattern, something that can of course be easily removed via a seperate LLM.
You seem to be assuming that they’ll be embedding hidden characters or something. That is not the case; the watermarking will be based around West specific words are chosen, meaning that if you type in those exact words, you’ll likely preserve the watermark. There are more details here: https://www.explainx.ai/blog/anthropic-claude-invisible-watermarks-c2pa-august-2026
I dont think, he’s not hiding any invisible stuff in ASCII, it’s not technically possible. But the pattern can itself be a watermark like the person before me just said
The difference between auto-generating hundreds to thousands of words, versus having to manually read and type all of those words out yourself is very obviously a lot more than “nothing”.
And that’s not even taking into account the fact that we don’t know what their method of watermarking is. From how they’re talking it sounds like it will be algorithmically embedded in the text itself somehow, through specific choices of words or something similar. If that’s the case, typing it out word for word will retain the watermark.
How tho? Will Claude be looking over our proverbial shoulders now, and if whatever we type just so happens to be word-for-word what it said, it can magically flag it?
You have a fundamental misunderstanding of how the watermarking works. There is no hidden metadata, but most likely, the statistical properties of the generated text are being changed.
If you met someone on Halloween with a face mask, and they’d always use the word cromulent in every sentence, you’d probably assume it’s your buddy Mark, who is about the only guy in your circle of friends who does that.
The model will produce a text where individual words at certain positions, or various n-grams encode a kind of fingerprint that will be an indicator for the text being processed by Claude.
Like, when people have, like, a specific accent or talk in a certain way, you can totally, like, figure out where they’re from, for sure.
So isn’t the solution then to just feed the Claude output through a second tiny LLM with the prompt to slightly rewrite the input? I mean you can probably use a 1b model to do that and that can run on practically anything nowadays
Yes, if the watermarking works like we all assume, significantly editing the output will likely break the watermark. They say this in their press release, as well as that the text must be a certain length to be watermarked. (Probably as a function of how the watermark works.)
I don’t think this is expected to be 100% foolproof. They certainly don’t make that claim; they explicitly say that the lack of a watermark isn’t conclusive evidence that the text wasn’t generated by their models.
I don’t think you will even need to type yourself. Tools removing the watermark will be available 3-4 days after this hits the market, fully able to be automated. Watermarks are not a solution that will prevent someone who wants to deceive you from doing so - only awareness that everything on the web is sus will.
Edit: The only alternative is an 100% digitally signature authenticated web, from the web login down to the posting. I am not sure this is a better way because it also means 100% transparency of everything someone does.
As I’ve said in response to a number of articles about this, all this means is that when someone wants to use Claude’s output for something, they just need to remember to manually type it in now, rather than just copy-paste it. Practically speaking, this changes nothing.
It’s some kind of statistical thing about what word is where, what character is there how often. Typing won’t help, but a rewrite will.
That was the author’s guess. Seems like this would only be detectable if it’s being used in enough text (possibly across multiple files).
What I’m saying is that if someone just types verbatim what Claude says, instead of just going “Ctrl-C, Ctrl-V”, that completely undermines this whole thing.
Their point is that Claude is now engineering its responses to specifically conform to a detectable pattern, something that can of course be easily removed via a seperate LLM.
You seem to be assuming that they’ll be embedding hidden characters or something. That is not the case; the watermarking will be based around West specific words are chosen, meaning that if you type in those exact words, you’ll likely preserve the watermark. There are more details here: https://www.explainx.ai/blog/anthropic-claude-invisible-watermarks-c2pa-august-2026
Very interesting read, thanks for sharing.
I dont think, he’s not hiding any invisible stuff in ASCII, it’s not technically possible. But the pattern can itself be a watermark like the person before me just said
The difference between auto-generating hundreds to thousands of words, versus having to manually read and type all of those words out yourself is very obviously a lot more than “nothing”.
And that’s not even taking into account the fact that we don’t know what their method of watermarking is. From how they’re talking it sounds like it will be algorithmically embedded in the text itself somehow, through specific choices of words or something similar. If that’s the case, typing it out word for word will retain the watermark.
How tho? Will Claude be looking over our proverbial shoulders now, and if whatever we type just so happens to be word-for-word what it said, it can magically flag it?
You have a fundamental misunderstanding of how the watermarking works. There is no hidden metadata, but most likely, the statistical properties of the generated text are being changed.
If you met someone on Halloween with a face mask, and they’d always use the word cromulent in every sentence, you’d probably assume it’s your buddy Mark, who is about the only guy in your circle of friends who does that.
The model will produce a text where individual words at certain positions, or various n-grams encode a kind of fingerprint that will be an indicator for the text being processed by Claude.
Like, when people have, like, a specific accent or talk in a certain way, you can totally, like, figure out where they’re from, for sure.
So isn’t the solution then to just feed the Claude output through a second tiny LLM with the prompt to slightly rewrite the input? I mean you can probably use a 1b model to do that and that can run on practically anything nowadays
Yes, if the watermarking works like we all assume, significantly editing the output will likely break the watermark. They say this in their press release, as well as that the text must be a certain length to be watermarked. (Probably as a function of how the watermark works.)
I don’t think this is expected to be 100% foolproof. They certainly don’t make that claim; they explicitly say that the lack of a watermark isn’t conclusive evidence that the text wasn’t generated by their models.
The words and the text itself are meaningless, it’s the content that matters
I don’t think you will even need to type yourself. Tools removing the watermark will be available 3-4 days after this hits the market, fully able to be automated. Watermarks are not a solution that will prevent someone who wants to deceive you from doing so - only awareness that everything on the web is sus will.
Edit: The only alternative is an 100% digitally signature authenticated web, from the web login down to the posting. I am not sure this is a better way because it also means 100% transparency of everything someone does.
There’s typically a keyboard shortcut to copy things sans formatting (plain text.)