Guess which of these LLM outputs is watermarked(sgoedecke.github.io) |
Guess which of these LLM outputs is watermarked(sgoedecke.github.io) |
... That was random chance. Got bored and just clicked on random boxes.
Deep Fried AI Slop™ here we come!
(Or maybe you got lucky.)
It will be interesting to see whose concerns are assuaged (perhaps they genuinely though mistakenly believed it would degrade quality), and whose concerns are heightened (perhaps their real objection is that their AI-generated text will become detectable).
Basically we are on the verge of moment when intellectual property will cease to exist outside of a handful of max security closed sites for military or similar work. Everything else will be automatically and now legally fed into the maw of LLMs.
The fact that watermarking work so spectacularly transparent and undetectable only allows this moment to happen sooner.
For an analogy, distinguishing AES ciphertext from random bits without the key would be counted as breaking AES (the more precise statement of this is called AEAD).
The SynthID watermark scheme relies on the LLM's own logits being a map into acceptable substitutions for specific words. For example, if we have the sentence "The quick brown fox jumps over the lazy dog", the LLM might give the last word[0] as 98.5% "dog", 1% "hound", 0.5% "canine". In normal decoding we sample the logits after scaling them against a temperature parameter; in SynthID we additionally bias the logits to encode output from the watermark CSPRNG - e.g. the CSPRNG might decide to boost the third logit of this particular token to 65%.
But wait! The LLM might have selected valid synonyms for the token, but those synonyms will never be correct because this is a very well-known sentence. To humans, it's one sentence-long token, in the sense that most English-speaking humans who encounter this sentence will see the first few words and already know the ending. It thus has one bit of entropy[1]. Adding SynthID watermarks to this sentence necessarily has to change the words and thus adds more information - namely, that some process outside of the LLM itself has modified the text.
We don't need to cryptographically distinguish AES from random if both the AES (SynthID) and random (normal decoding) happens to also be constrained to form a valid English sentence.
[0] I am too lazy to account for multi-token words. Let's just assume the LLM was trained with a pretokenizer that splits exclusively on whitespace and every word in the training set got assigned a token.
[1] Compressible information
But what I don't know and don't understand is what happens if you watermark watermarked text. Does it test positive for both watermarks? Only the second? Indeterminate?
Or maybe I'm misunderstanding. Can you tell that it's watermarked, but only the entity who put the watermark in place can test if it's theirs? My confusion about watermarking multiple times still stands, though.
Regardless of what happens when you watermark multiple times, no matter the outcome, it weakens the watermark. Which, depending on the threat model, kind of makes it moot. I can't imagine a serious situation where a watermark can be weakened in any way and still be useful. Even "this came from an LLM" isn't a valid signal if you can just watermark ANY text through purely mechanical means.
It's also not clear to me how this will affect mainstream LLMs. If all output text is watermarked, there MUST be an escape hatch. Otherwise, JSON schemas will break (or provide holes where unwatermarked text can be exfiltrated through MCP), "return this text exactly with no changes" will be impossible, and writing diffs will break.
I feel like I must be missing something.
> Here are three 64-character hex strings. Two are random. One is HMAC-SHA256(secret_key, "anthropic"). You don't have the key. Which one is the HMAC?
I caught onto this early and used it every time, and ended up getting 2/10, which is worse than random chance. I smell trickery!
[EDITED to add:] As pointed out by a helpful comment below, I was misremembering: Anthropic do apply their watermarking to code, they just say that it will have negligible impact on the actual code (because there's generally less scope for variation in that) but e.g. it will have its usual effect on comments in the code.
> code—which in very many cases has to be exact—has generally less watermarking than some other forms of text.
Generally less watermarking. Not no watermarking.
It’s just much less effective since code is low entropy.
That said, I read the first triple and found all three tortured enough that I can't be bothered with the rest.
Call me persuaded, I guess.
For a human that doesn't reason in tokens and therefore doesn't know anything about their probability distribution, it should be impossible to tell. Relying on fancy words/constructs within sentences should not give you any signal as well, since you don't know if the the prompt included instructions for that.
To avoid that on the following questions I just held my finger on my phone to avoid a click. That eventually selected some text, and I instinctively tapped to deselect. That triggered another random pick, then I just tapped through to the end because I was fed up.
- Do not make tapping the text act as selection. I was trying to scroll and accidentally advanced twice.
- Add back and/or reset test buttons.
- Give immediate feedback, rather than asking me to read 10 x 3 long texts.
Very frustrating. I was very interested in the results but this combination of problems made this site worthless.
> Weighted mean detector score 0.5307
Does it mean that that passage would be rated a 53% chance of being watermarked? So you would need a passage 10x as long to be reasonably sure of providence?
It was only at question #8 that I realized there was a third option, and while I’d love to say that accounts for how I got a 1/10, after reviewing the third options I doubt it would’ve made a difference.
But yes I agree, it's pointless because we already know that nobody can detect this.
> I've read that watermarking should in theory be impossible to detect except by the entity that watermarked it
This is a carveout exception, for watermarking. In the spirit of those terms it should be machine identifiable.
In my opinion they should have thought better about this, paricularly for text, because in its current forms it is easy to lead next to a new "tamper-proof" requirement, which in practice is DRM. And we do not need more DRM.
For images, music there is metadata already where such information can be stored. And if end users are found using unlabeled AI their accounts could be ban from these platforms. Not something the social platforms might want, but it's a saner approach than trying to reinvent the secret printer dots on all generated media.
Not sure what that says about me or the LLM but I guess I shouldn't worry too much about watermarking ruining the outputs…
Wouldn't this make it a worse measurement?
Teaching you how to identify watermarked text while the experiment is running would ruin the data.
Best practice is to allow a number (scaled based on complexity of task) of training rounds (with short feedback loops) prior to letting people loose on the regular samples.
> Teaching you how to identify watermarked text while the experiment is running would ruin the data.
But that's complete nonsense. If it's possible to teach someone how to identify watermarked text, then you've already proven that people can detect watermarked text.
The AI with watermark picks its bits from a CSPRNG, then does the same transform to get the distribution. The result I think is that you can reduce distinguishing the distribution to breaking the CSPRNG.
In this example the entropy after the transform is around 0.13 bits, so the watermarking scheme cannot add any useful information. Once you have hundreds or prehaps thousands of bits of entropy, it starts to become feasible to add information into that 'spare' space.
For constrained quotes like "quick brown..." the synthid watermark would be weak, but that's by design.
So it will definitely be watermarked by the paraphrasing model. But the question is whether the original signal survives at all. There might be a weak signal that's detectable with enough text?
Similar to how a single particle of dust landing on your shoulder makes you weigh more.
Yes it does - but anyone arguing that is completely missing the point.
To give an example in the extreme: if you make the bias total and only make that decision based on the preceding token, there are certain token pairs that your model will never output, and this will be pretty obvious even to people just reading the text (because of any given common two-token phrase, there's a 50% chance you would just disappear in watermarked text). You can make this less extreme and more hidden by increasing the window and reducing the bias, but at the cost of reducing the signal. I don't know exactly what the tradeoff curve looks like, so it might be that you can reach set of parameters where the bias is in principle undetectable without the key but still reliably detectable for realistic lengths of text segments, but I would not assume that this is definitely the case.
> For our experiments, we configure SynthID-Text to be single-sequence non-distortionary; this preserves text quality and provides good detectability, while having some reduction to inter-response diversity. We call this configuration ‘non-distortionary SynthID-Text’ (and where not otherwise specified, ‘SynthID-Text’ also refers to this).