Spymarks, Not Watermarks(brand.io) |
Spymarks, Not Watermarks(brand.io) |
If it is text, copying text alone and not the file will it not remove it? Massage the text with Ai and vola spymark gone, don't you think?
No. The mark is hidden in the word choices.
See the demo in the article.
It's not "your work" it's the bloody AI's work! That's the whole point.
> > Because I always wanted to coin something. Please don't forget me.
Watermarks are not "spymarks". They're DRM. I wouldn't worry about advertisers tracking conversions. I would worry about the "analog hole" being closed. Think of no longer being able to even photograph your phone screen, because pixels on it carry digital watermark that's robust enough to survive being photographed - I.e. the kind currently used to tag AI generated images - and then every phone and computer refusing to display resulting photo because the app disallowed capturing its pixels.
> On that note, one way we can prevent it is to assert that all our content is byte-for-byte identical with the last known trusted stage of what we have produced
That doesn't help with things like the typical use of SynthID where the spymarking is done by the same process generating the content, so there is never a clean comparator. (It also wouldn't be useful anytime it is inplemented as part of a transformation—compression, etc. —step, for the same reason.)
This has been going on for a while with Facebook. They seem to embed custom metadata tags so that images shared outside the platform can be traced back:
https://stackoverflow.com/questions/31120222/iptc-metadata-a...
Stop using links instead of words. Your comment is literally unreadable without going on to other websites.
No, the words could contain steganography. Use links to be safe!
:D
What will be the next? I will be unable to see the domain of a link on hover/longtap and have to trust random links like on a search engine?
MUTINY against hn!
Or... add some noise. Just align the last bit of every pixel channel with a random bit sequence - and Bob's your uncle.
First the low-end laptops and phones (and probably later, most of them) will incorporate some low-level driver that is constantly scanning for these and passing them to a helper app to phone home. I assume this is something Apple will, to their credit, refuse to do[1] but I don't think other OEMs will have any qualms based on what they already do with their TVs.
[1] (though they don't do this kind of thing out of altruism, but because their cash cow is app store rents and fat hardware margins, not third-party advertising.)
I still use IRC on the daily, but someone mentions Discord at least once a month on there, and sometimes tries to whisk people away to that side. There are also people hooking up LLMs to IRC bots and joining them to channels without permission, then when you complain or kick/ban their bot you're somehow treated as the rude one. It's very hard to entirely get away from all the crap anymore.
So I’m not writing off tech as a whole just the adtech companies being a lost cause.
It's a shame however, how low quality and vibecoded the live examples are. The first example says "Toy example; not SynthID.", the second one is a generic spectrogram and the third one has an identification space too small to be useful (173 in decimal). I was hoping to see more realistic scenarios to learn how these new watermarks are being applied, instead of generic steganography.
Like it choose between "winding" and "curving" but there are many uses of curving that probably can't be replaced with "winding" like "her gently curving thighs" with "her gently winding thighs"
But I'm sure there are some intricacies I don't understand. Anyway, very cool website, thanks for sharing it~!
TrojanStego: https://arxiv.org/abs/2505.20118 Improvement: https://arxiv.org/abs/2606.09411
The "tracking" bit is kind of nefarious, but that can be removed as a concern if the thing that is being tracked is agents, not users.
Digital Rights Management isn't just about restricting what you can do with the content, which is often futile anyway. Another way companies can manage their rights is by making sure pirates are properly identified and caught.
You can’t definitively prove the absence of a watermark. You can only prove the watermark is there. Once you do prove it’s there, the thing that carries the watermark changes in some way — it is “burned” or tainted?
There must be value in having a visible vs an invisible watermark, or in declaring that a work is watermarked without revealing the hidden mark, or having two marks — one that is publicly verifiable and another that is hidden?
If the process itself can be defeated through adding entropy (or more generally by revealing the watermark algorithm) then is that not security through obscurity, which is to say it is a one-shot rather than a general system that is doomed to become obsolete over time?
Something feels off about a technology based on being hidden but whose only value is in being revealed but I feel dumb for not being able to be more specific about what feels wrong! It could simply be that anyone who can verify the presence of the watermark also now has a tool to tell them when they’ve successfully scrubbed the watermark off the work, so the verify tool has to be kept secret which in turn limits its usefulness.
I spent half a day messing around with it and I was very impressed by how robust it is. I couldn't get OpenAI to stop detecting their own SynthID without completely trashing the image.
I don't think its fair to say that metadata on apps will be safely removable in the future.
https://www.abc.net.au/news/2026-08-18/what-happens-if-prope...
How do you spymark text that someone else wrote? You can't change the words or they'd notice
Text the whistleblower only reports on, well, if they got it from a computer system, there's already precedent of altering word choices, typos and punctuation in e-mails and memos to create unique per-recipient or per-recipient-group versions, which allows companies to trace leaked transcripts reported by press back to source of the leak.
Might be technically more accurate, but it’s an inscrutable name with no chance of proliferation beyond technical people. If the goal is to rally people to your position, you need a name people can identify.
Of course, the absence of a watermark/spymark doesn't prove that the source wasn't AI generated. But the absence provides evidence that it wasn't.
AI companies should simply use watermarks in the responsible sense: they should indicate that the material was AI generated, not include personal information in it.
Tech like SynthID I see a net positive especially since it doesn't degrade text quality. I dream about a browser extension running at all times that makes text more translucent based on the confidence of LLM writing[0].
This article's suggestion of using it to unmask whistleblowers is very interesting and not something I'd thought about though. Still not convinced that spymark is a better name though.
[0]: Sean Goedecke's Deckard is close but it would rather invisible than bright red https://www.seangoedecke.com/deckard/
When some processing is required by eu or eu member state law, that processing doesn't need explicit consent.
One could also argue the watermark tracks the generated content, not the person
We also recently had this with LG spy-TVs. Cars here in the EU also spy on people, allegedly to show how alert they are. Perhaps they sneakily upload that information somewhere ... Facebook also has the spy-glasses now. People getting angry about Flock-spy-cameras.
It seems we are now in the age of spying of everyone at all times. Future spying will be done via even smaller devices.
The concern is valid, but microphones are far worse. They're simpler, smaller, extremely sensitive to sound and an order of magnitude cheaper, both the mic itself as well as any spying with it.
It is possible to record voice using few bytes, to send later. It's further possible to transcribe cheaply into text, and analyze said text.
And mics are already everywhere, including in devices that do not need them, as well as speakers that can be rewired by software to act as microphones.
Because they have this pathway only they can access, their key and signature can be trusted. I’m not gonna trust Joe Schmoe’s signature that “No I didn’t use AI” unless I already trust Joe Schmoe (and in which case, he doesn’t need a watermark, I’ll just believe him when he says it).
It’s your choice if you believe them or not, I like Apple and I wouldn’t use that feature.
The pretending that this is the same thing, that Apple is sneaking something past you when they’re showing you that they’re trying to do it right is a bad faith argument.
I think you may have an outdated view here
But fuck it, just brand all our brains with "SLA Industries".(fictional dystopian corporation ruling future)
The article calls out watermarks intended to deter counterfeiting as explicitly being not spymarks. Watermarks can tell you about the items marked, not about the person using/creating it.
Bank notes are mainly protected by things that are hard to create without very specialized machines. You can't rely on anything hidden staying unknown and once you know a steganography scheme you can also control it.
> Tech like SynthID I see a net positive especially since it doesn't degrade text quality. I dream about a browser extension running at all times that makes text more translucent based on the confidence of LLM writing[0].
False confidence is a lot worse than no confidence. If such an extension ever becomes popular, people will take anything not marked as AI as gospel.
I'm not convinced spymark is better than just "invisible watermarks", spymark to my ears sounds designed to be sound very negative when invisible watermarks are not always negative, e.g. the counterfeit bank note example.
The article spent quite a bit of energy explaining why the word choice, seems like you're just ignoring that? Also, bank note watermarks are not invisible. Watermarks are not invisible, as the article (again) took pains to explain. Tech like SynthID I see a net positive especially since it doesn't degrade text quality.
It absolutely does. It constrains high-entropy word choice so it can "store" other things in your text. Your text actually has information removed from it. I dream about a browser extension running at all times that makes text more translucent based on the confidence of LLM writing.
Sounds very much not worth the anti-consumer, anti-privacy aspects which (again) the article explains.One is proof of authenticity, other is tracking tool.
I wonder if we could have a real, open and direct democracy. All the models we have right now work via indirect clowns. Then again, looking at how some people vote, perhaps direct democracy can only work if people are clever.
To do politics right take skills, and I don't expect the average mechanic to be better at it than the average politician is at fixing cars. How should I know if we should subsidize organic farming, ban alcohol sale after 8PM, or increase the defense budget? At least in theory, politicians are professionals who deal with these kinds of questions, they are supposed to know the technical and social implications, or find experts to help them if they don't. Some people think they know, and judging by how stupid most of their ideas are, they don't. I don't blame them, it is just not their field, and my ideas are probably just as stupid anyways.
The problem with these spymarks is that they can be used to include data that's completely invisible to users - even potentially to sophisticated users and the programs that consume the marked files. So while I can make an informed decision whether to share a picture, I may not be informed about any spymarks. Vs a normal watermark that aren't designed to be invisible.
The closest thing I found to "defeating" SynthID was to put in a normal photograph and ask ChatGPT to make some utterly trivial edit, and then the output got flagged with SynthID even though it is essentially an unmodified photograph.
I would. Politicians are pre-selected for people who want to lead and that's the last kind person that should be allowed to.
(Edit: hmm, without the "https://" it seems to depend on the browsers ability to recognise a URL.)
"for stenography (link)"
Or to use another HNism
"Stenography[1]"
Those interested could click it, those not could still read the comment.
Even if you trust them, maybe as an user you can be ok with that. As a non-user who will talk with people wearing Apple Watches, I disagree being recorded and my conversations with the watch owner summarized.
Where do I disagree for that ?
e.g. Johnny's accused of something white collar, did he ever make any prompts that suggest how early on he was aware of {X} and further indicate how he moved to frame it?
That's a requirement that varies by country.
They’re asserting that all the audio is done on device, and the results of encrypted so they can’t access them even from the backups.
Unlike… EVERY… other tech company, it is in Apple’s interest to be privacy-focused.
Even if you just have to believe them, which you do pretty much, they’re the biggest name pushing for privacy in the world right now. They make more on selling devices than they make on ads and behaviors. It’s in their interest to not lie.
Of course, at that point it depends on how much you trust their PCC.
But make no mistake, Apple itself is an ad company and that sets all the incentives that matter.
Oh, really?