AI companies leak data to advertisers [pdf](jorgegarciaherrero.com) |
AI companies leak data to advertisers [pdf](jorgegarciaherrero.com) |
We have all become Milhouse now.
Not good at all.
Perplexity does this. Visiting a past perplexity search url exposes your full conversation.
I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link
Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped
There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines
This is shockingly lax approach to data security and privacy by design
It’s not like someone’s gonna guess that URL… right?
The amount of sensitive information that accidentally gets left on screenshots is pretty large. This is a pretty massive security issue
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).
If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).
I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.
>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.
Who thinks OpenAI or any big company for that matter give a crap about them? This isn't a popular sentiment at all, it's just patently false.
Access through: https://ai.ivx.run/chat/
First: You don't want to leak information about your users to advertising networks because it's going to leak, get back to your customers, they're going to figure out you're doing it and get really angry.
But second and more importantly - it's a much better business model to collect that data for yourself, keep it in house and then you control how you use that data to target ads which gives you a massive competitive advantage in selling ads because you have unique targeting data.
The way meta does this now is the model, they don't give the advertiser a list of the people you're going to show the advert to, the advertiser gives you a list of characteristics they want to hit and meta decides who those people are.
Apparently they have not learned, or they learned the wrong lesson. I do not think they leak it intentionally as hoarding is typically more profitable than selling but anyone who has been in the industry for some time knows that the move fast and break things attitude has caused enormous amounts of data leaks.
What happened? Oh no!
How terrible! That’s just, that’s just awful!
How terrible! Oh no!
Someone should have to investigate, but I suppose it's all "legal"?
In EU law it is very likely against GDPR.
Screenshots of conversations. TikTok received screenshots of Grok chats during sharing, exposing the actual visible conversation content.
Conversation-derived content tied to persistent identifiers, including prompts and automatically generated chat titles revealing sensitive facts. "Salary 85k NYC: mortgage 280–350k".
I also accidentally paste random stuff into input boxes all the time.