136 comments
I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.
This partial prompt data might potentially be used to "pre-warm" some kind of cache.
But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.
OpenAI’s statements in response to the Millenium Prize (and related) disputes I think are a pretty obvious example of this in practice. One man’s “user prompts” is not another’s “reasoning trace scratchpad”.
This comment by Falserum on the mathematics research post articulates it well:
I'm pretty sure it is used for this; but rather than for anything nefarious, my guess is that this info is then fed to a classifier model to ensure that users of ChatGPT-the-service (as opposed to the OpenAI inference API) are actual humans, rather than agents trying to circumvent having to pay API pricing.
Sounds good.
Perplexity does this. Visiting a past perplexity search url exposes your full conversation.
I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link
Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped
There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines
This is shockingly lax approach to data security and privacy by design
If you click "provide a shareable link" you should decide (and behave) as though that made it public.
I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."
It’s not like someone’s gonna guess that URL… right?
I also accidentally paste random stuff into input boxes all the time.
however - agree that this is not great - espeically if chat TTL is long. someone who gets your URL can read everything you're asking (eg. by sniffing your network/accessing your browser history)
We have all become Milhouse now.
People love inequality when it's a celebrity they adore living large. Someone who hasn't scammed them and has demonstrably improved their life and the lives of others in a tangible way. The deeper the con, like Trump and Elon, the more damage to trust.
Unless you become a target of the government. Then people a lot worse than any teacher you ever had will be looking through it, and they, unfortunately, do care.
My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.
Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).
If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).
I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.
Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.
>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.
Read the full thread on Hacker News →
Related stories
- The Verge · 0 points · 2 days ago
- Can you forget how you feel about Meta?theverge.comThe Verge · 0 points · 9 days ago
- The Verge · 0 points · 4 days ago
- Can John Ternus find Apple’s next big thing?theverge.comThe Verge · 0 points · 9 days ago
- Hacker News · 73 points · 8 days ago
- The Verge · 0 points · 11 days ago