136 comments

pbasista1 day ago
Tangential:

I have recently noticed that e.g. ChatGPT, when used from a web browser, periodically sends unfinished prompts to their servers, namely to the `conversation/prepare` endpoint, without waiting for the user to actually send it.

This partial prompt data might potentially be used to "pre-warm" some kind of cache.

But it may also be used to track the user's writing cadence, error correction style and evolution of their stub ideas as they are being formulated into a prompt. I would assume that such data could also be sold to the advertisers.

ShinyLeftPad1 day ago
I wouldn't be surprised if their privacy policy would say "what you send is private" and then it wouldn't apply to unfinished prompts on technicality
jwstillwater1 day ago
This is my concern as well- the same wiggle-room methodology that allowed a business to claim not to “sell or share” PII, because “user data collaboration” was not part of the legal definition prior to CCPA.

OpenAI’s statements in response to the Millenium Prize (and related) disputes I think are a pretty obvious example of this in practice. One man’s “user prompts” is not another’s “reasoning trace scratchpad”.

This comment by Falserum on the mathematics research post articulates it well:

https://news.ycombinator.com/item?id=49649992

derefr1 day ago
> But it may also be used to track the user's writing cadence, error correction style

I'm pretty sure it is used for this; but rather than for anything nefarious, my guess is that this info is then fed to a classifier model to ensure that users of ChatGPT-the-service (as opposed to the OpenAI inference API) are actual humans, rather than agents trying to circumvent having to pay API pricing.

kridsdale11 day ago
Enough Meta executives have been hired there. Expect the same behavior over time.
sroussey1 day ago
You open the same convo on another device and the partially written text is there to continue. Does it not do that for you?
port11about 13 hours ago
“Oh, we’ll store your unfinished thoughts, privately typed into a text box, on the off-chance that you want to continue writing it on another device. And no, we won’t ask you nor give you any semblance of true privacy.”

Sounds good.

kdaniel_031 day ago
It's the same lesson as the Navier-Stokes credit fight earlier this month. Buckmaster and Alpoge had their unpublished drafts in private Codex sessions and OpenAI says nobody saw them but admits de-identified product data may have improved its models. There it's training data, here it's ad trackers. Either way, prompts and results that should stay private don't. Thats why even though open models aren't perfect it has to win. You can skip the app and run the model yourself.
postalcoder1 day ago
My least favorite trend I’ve noticed with so many AI chat services is they seem to equate a UUID in the url with privacy.

Perplexity does this. Visiting a past perplexity search url exposes your full conversation.

albert_e1 day ago
Security by obscurity -- such an age old anti-pattern!

I believe many AI tools like Gemini generate publicly accessible URLs when we click "Share" on any chat conversation -- and expect users to then own the lifecycle of that link

Depending on how the link gets handled -- by the browser, device OS, any hooks/plugins/extensions, aggressive telemetry, social media url previews, preload/prefetch, wrapping and url shortening, etc as it reaches the intended user -- there are countless ways in which the URL can be indexed and scraped

There was a issue not long ago when Claude artifacts were indexed en-masse by Google and other search engines

This is shockingly lax approach to data security and privacy by design

kevindamm1 day ago
The same assumptions are true about giving any human that shareable link. They could pass it on to anyone, screenshot it, paste it into their own session. This has been true since before "share with link" permissions on Docs and elsewhere.

If you click "provide a shareable link" you should decide (and behave) as though that made it public.

I'm not saying it's good privacy posture on the side of the companies, but how else do you think that would work if there isn't any authentication step for the person viewing it? Even with authentication, "three may keep a secret, if two of them are dead."

postalcoder1 day ago
Chat UIs are a minefield of “if you accidentally click this your data will be shared or trained without you realizing it!”
40fourabout 23 hours ago
It certainly doesn’t expose it to anyone else besides you (when you are logged in), unless you explicitly select the share option.
msdz1 day ago
Genuinely asking: If you don’t share the UUID-based URL yourself, what makes it not privacy-friendly?

It’s not like someone’s gonna guess that URL… right?

postalcoder1 day ago
Yes, technically, guessing a url is impossible. But browser histories are stored in cleartext and trivially accessible to sketchy actors.

I also accidentally paste random stuff into input boxes all the time.

someonebaggy1 day ago
Isn't that equivalent to a password? Knowing my password exposes my full data.
layerv-ai1 day ago
not as bad since the blast radius is only your chat vs. knowing your password exposes all your data.

however - agree that this is not great - espeically if chat TTL is long. someone who gets your URL can read everything you're asking (eg. by sniffing your network/accessing your browser history)

wtetzner1 day ago
You don't store your password in the URL.
In an old Simpsons episode Lisa gets to visit the Teachers room, where all the staff are making fun of the children. Groundskeeper Willie is pantomiming Milhouse “Oh I am Milhouse, I tell all my secrets to Willie since I have no friends!” and the teachers laugh. Later something embarrassing happens to Milhouse and he immediately runs away crying “I have to tell this to Willie!”.

We have all become Milhouse now.

Avicebron1 day ago
Inequality has eroded trust in society in ~50 years, a lot of the old models (heh) of how we see the world aren't relevant. It's hard to exist when everything around us is adversarial.
Data brokers were around 50 years ago. They just had more limited sources to draw from.
gmd631 day ago
It's not inequality. It's who we've chosen to reward. Adversarial people have eaten a lot of the world and that's because we let them.

People love inequality when it's a celebrity they adore living large. Someone who hasn't scammed them and has demonstrably improved their life and the lives of others in a tangible way. The deeper the con, like Trump and Elon, the more damage to trust.

keybored1 day ago
It’s just inequality full stop. Trust is a word by wannabe-Bernays polsci people who would just so dearly want the proles to trust their betters, but for some dastardly material reasons they don’t.
This can't be real I am go gonna ask Willie
davsti41 day ago
But, Willie isn't real?
consensus11 day ago
But we are not really like Milhouse. In reality the teachers didn't care about us enough to make fun of us in the teachers lounge. We were just that year's batch of work soon to be forgotten when they move on to the next. This is exactly the same as the advertising companies. You are just a member of the cohort of a particular target campaign. Nobody cares about you or even knows you are part of that target cohort.

Unless you become a target of the government. Then people a lot worse than any teacher you ever had will be looking through it, and they, unfortunately, do care.

j4k0bfr1 day ago
This is a bit surprising to me, considering how much AI companies love to hoard data. Especially since some of these ad companies are direct competitors!

My best guess is that these ad mechanisms are a bit rushed and/or that investor demands for profitability are fighting against company self-interest.

Edit: I guess some data will always need to be leaked for AI chat ads to be most effective. But I imagine AI companies would rather deliver the targeted ads themselves rather than letting competitors do it for them. It would be scary to see AI companies become ad companies too (instead of just hosting them).

amarcheschi1 day ago
I'm taking an onboarding process for an Ai company helping other (much) bigger Ai companies and the amount of vibecoded platforms and documentation is staggering. Like, training process so broken that the platform just doesn't load sometimes, things that have never even been tried are published and you have to use them and they suck so much because it is apparent that no human ever touched that and probably wouldn't want to
Forgeties791 day ago
Doesn’t sound like your new company is going to be very helpful for other AI companies lol
alansaber1 day ago
AI companies rushing an implementation? Surely not :).
Forgeties791 day ago
No no it’s hyperscaling
mrweasel1 day ago
So I only read the abstract, but the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.

If that's the case, then I'm not surprised at all. Actually I also wouldn't be surprised if they sold the data, but that's a different story. If we look at OpenAI for instance, they have on multiple occasion shown that they do not have the operational experience or resources to run their services in a safe and secure manor, nor do they frankly have an impressive up reliability (in terms of operational stability).

I'd support your guess that all of this is rushed in an attempt to push for profitabilitet/growth.

kjs31 day ago
the question is if the data is leaked, by accident, or if it's deliberately provided. My guess is that we're talking about the first scenario, and that this is an accident.

Read the T&Cs. If there's even the tiniest bit of "we might provide your data to third parties for the purposes of...", it's not an accident. Virtually all the AI company T&Cs I've looked at had weasel words that open that door, because it's obvious to anyone paying attention that cramming ads into AI products is the next frontier in AI revenue streams.

duskdozer1 day ago
Well, some data is deliberately provided at least. It's not a situation of other apps managing to grab the data like the facebook-localhost exploit

>The most prevalent third-party services included in CSP headers belong to Google Tag Manager (googletagmanager.com), Google Analytics (google-analytics.com), and Google Ads (googleadservices.com and doubleclick.net). Yet, as Table 8 shows, CSP policies commonly include other prominent actors in the advertising industry, such as TikTok and Meta.

dgellow1 day ago
I don’t think it matters much if it is intentional or not, no? In both cases it’s a breach of privacy that should be punished. Though I myself do not believe they are selling, I don’t think they reached that level of sophistication yet, they seem to be way too hacky for that type of scheme. But that will for sure come later on

Read the full thread on Hacker News →

Related stories