Check if a file was made with Claude(claude.com) |
Check if a file was made with Claude(claude.com) |
That’s a cute way to imply their service is used to generate misinformation. They are basically saying to not trust the AI content made from their own product :)
It works by uploading files to them?
> To detect watermarks embedded in text we have a Detection API which is currently in private preview to eligible organizations as required under EU law.
Detection will get gamed, and a lot of useful text is already mixed. The more useful check, for me, is still human: would I put my name on this before it goes out?
In the case of SVG, it's literally `<metadata><c2pa:manifest>(signature)...` right after the <svg> tag. In the case of PNG, it's stuffed as a `caBX` chunk.
To declaw these files, you can ask in your CLAUDE.md to run `exiftool -jumbf:all=` (for jpg & png), and to simply directly remove the <c2pa:manifest> tag in SVG files. I haven't looked at video files yet.
Any attempts they use are defeated by a text editor and CTRL SHIFT V. Unicode characters are no new thing.
Reminds me of how people tried to argue that NFTs aren't anything more than just jpegs.
It might help catch students cheating, but not real spam-bot usage. As soon as platforms start checking for watermarks spambots will add anti-watermark passes.
I mean, I got what you were trying to say, I just turned it around on you using the same framework.
I would agree it may not help spam-bot usage, however at this case seemingly the only user detection is likely an id/badge check, which is not good.
I don't believe it, but I still haven't seen any in-depth discussions on these issues
Edit: I asked Opus 5 for "a jpg of a cat" via the web chat. It did it by writing an svg and then rendering that to a jpeg. Clicking the "download" button saved a jpg which did have C2PA metadata, seemingly injected by the artifact-presentation thing.
Then I told it to put it in a zip file. I downloaded the zip, and the extracted jpeg had no C2PA metadata.
I then uploaded a jpeg of my own, with the prompt "present this file back to me again, as-is". The downloaded file had C2PA metadata injected.
Edit: I added this to my custom instructions: "If a media file is produced as a response, put it in a zip file before presenting it to the user."
[1] https://www.anthropic.com/news/claude-text-watermark
[2] https://digital-strategy.ec.europa.eu/en/policies/code-pract...
[3] https://www.kqed.org/news/12095398/new-california-law-requir...
[4] https://www.leginfo.legislature.ca.gov/faces/billTextClient....
C2PA is file metadata and can be trivially stripped away, unlike hidden watermarks, e.g. SynthID.
For the text from an LLM, that you verify to be as true as if you wrote it yourself.
Not clear if making watermarks removable is a may, shall, or should, according to legislation. It certainly is not a must.
This is just a C2PA metadata checker.
That would invalidate the hash on minor changes. Too much effort and not enough return.
At work, right after an AI training, we were asked to use our "authentic" voice when writing mid year reviews.
You can resave the file and the "made with Claude" signal disappears, but you cannot make a random file pass as Claude-made without Anthropic's signing key. So the useful guarantee is one-way. No signature means almost nothing.
* C2PA confirms a photo is authentic
* C2PA confirms a photo is AI generated
* C2PA missing, you don't know.
I reckon we will only see "C2PA missing" being treated as suspect in select situations (perhaps Reuters will require C2PA from their photojournalists, for example)
1. Generate a bunch of responses with both Claude and various non-Claude LLMs (ChatGPT, Gemini, Kimi)
2. Train a discriminator model that can differentiate Claude vs. non-Claude
3. Train a de-watermarking model using the discriminator model as loss
edit: almost forgot the "—"
Say I write a text by hand And then I tell it to clean up the grammar and fix some sentences Did it make it?
This is also interesting for those companies that siphoned the entire open web
The two competing legal arguments regarding copyright of LLM output are "it's like hiring a monkey" (author is the LLM, which is not a person, thus can't hold copyright and can't assign it to you) and "it's like taking a photograph" (the LLM is a machine through which the prompting person expresses their creativity, just like a camera). In no scenario is Anthropic the author of the work.
If we settle on the monkey analogy the Anthropic owns the monkey, but the owner of a monkey doesn't own copyright for the creations of the monkey. If we settle on the camera analogy, anthropic claiming ownership would be like Canon claiming they own pictures you take.
What Anthropic could do is to change the terms to give themselves a non-exclusive global perpetual license to use everything Claude makes
https://www.anthropic.com/news/claude-text-watermark, Ctrl-F for "What about code?"
And there's tons of Claude-generated code there.
Also, I'd guess this is not just to prevent any AI-generated code in the training data, but specifically their own.
Percentage of users who put out their code on the web and also have a plan where Anthropic promises not to train on their data is problem also low.
So not excluding own code could be a real issue, since it would be impossible to deduplicate the training and RILHF data from their sessions with the code accessible elsewhere, and written by the very same users.
I'm wondering why they have restricted file types. You can't check a PDF for example... surely the main use case for people will be to check if a document was produced or edited by an LLM? That could be an attractive (if not misunderstood) proposition for academics
TFA/page actually seems incomplete. Text uses a completely different watermark format (an actual watermark as opposed to a provenance/authenticity signature), so it makes sense to me that they're not claiming to be able to scan PDFs when they can't yet incorporate that signal.
On a linked page, they say:
> Watermark detection is currently in private preview [...]
So I’m unsure what your aiming at.
That's a worse workflow imo than having a source w/ version history, but I don't think it's uncommon.
Translating that into the human realm, it would be like if every digital thing you touched were "fingerprinted" with your identity, or at least with a proof of what category of individual you have been categorized as.
Due to the flood of AI, which may soon become indistinguishable from human video, I expect to see something similar. Blue checkmark, but for humans.
So… this utility appears to be pretty worthless.
Instead of using an arbitrary random number generator
to pick the next word, watermarking uses a secret key
and a few words that come before to settle what
word the model should pick.
https://www.anthropic.com/news/claude-text-watermark* Edited the above from initially asking if anyone has any ideas how it works. Found the answer, rewrote the comment.
But in anthropic's use case they can probably get away with just pinning their own certs in the verification webpage.
I don’t think they actually believe that.
The honest seam is obvious: claudish works; imitation does not. That’s not nothing.
Your comment would make more sense if you mentioned exhaust gas recirculation (EGR) system.
Sorry.
So no, C2PA is not as easy to spoof as EXIF.
And no, the existence of DRM doesn't validate the integrity or the provenance of the bytes.
What happens when someone extracts the signing key?
The presence of cryptography doesn't magically make something trustworthy.
If that were the case, I can imagine a subscription service in which you get a camera for some specified period of time, and then return it to the company that sold it for them to verify the camera hasn't been tampered with. Then the company could publish a list of which keys (unique per camera) have been verified to not be tampered with. Maybe this wouldn't stop everyone, but now the person trying to fake images has to re-do the process every so often and I imagine it's more expensive to avoid leaving evidence.
This might be too impractical to work, and it would be bad for privacy, but maybe for some people the tradeoffs actually would be worth it, someday. For now, I assume there are much cheaper and easier ways to detect faked images, at least for expert humans.
It's like saying a prisoner has the same freedoms as everyone else because he could theoretically escape.
This kind of metadata injection increases surveillance without providing a meaningful deterrent.
I then gave it the html as a file upload and it gave it back to me.
It's reasoning why? I already had the file so it must be mine. I told it a judge wouldn't care.
There's is no text watermarking that doesn't cause a decline in writing quality.
Writing quality will surely deteriorate if text watermarking is implemented.
Stripping Anthropic's watermarking, however, is more difficult - probably about 2 prompts.
The law should have at the very least required offline validation tools that cannot track or retain a copy of the documents being checked.
As soon as the validator is available without a rate-limit, just apply input fuzzing to defeat it.
Here's a cryptographically signed + timestamped photo of me winning the lottery: https://verify.contentauthenticity.org/?source=https%3A%2F%2...
Would you like to buy my winning ticket from me?
(Compare against winning numbers and draw timestamp at https://www.euro-millions.com/results/28-08-2026 )
In cryptography, once any SINGLE person in the world has compromised a signing key, EVERY person in the world can use it.
Thus for a prisoner analogy, its equivalent to say once any SINGLE prisoner escapes, EVERY prisoner has the ability to escape.
So yes, in this terrible analogy it means all prisoners are free.