Lasso Research tested SynthID-Text watermarking across six models and found it changes tool-call correctness and weakens refusal under prompt injection. On some models, watermark-induced behavioral churn exceeds what a…

58 points•nisosguy•5 days ago•72 comments•

72 comments

WithinReason4 days ago
This is getting tiring. Watermarking has no effect on model output quality when implemented correctly. It's somewhat like swapping a random RNG seed to the seed 42, and detecting what the seed was from a random sequence. The sequence generated from the seed 42 is just as random as any other seed. There couldn't be a quality difference. And yes, the output from an LLM is a conditional random sequence of tokens from a distribution determined by a model.
Phemist4 days ago
The article has a pretty decent summary of the watermarking algo though. This reads as a pretty dogmatic statement in comparison.

In your analogy: What if seed 42 specifically causes poor quality behaviour (in some contexts specifically). Normally, these quality differences will be washed out because the seed is random, now it is no longer random, so shouldnt we check into specific behaviour under this specific seed?

Phemist4 days ago
Refusal behaviour specifically is interesting, because if you can point out specific cases where refusal behaviour significantly deteriorates due to the watermarking, it may create a token "route" that may be possible to exploit by adverserial prompters. Static hazardous prompt refusal belies the fact that actual adversarial prompters will adapt their techniques iteratively and gain way higher compliance rates.

My idea would be that the ngram size over which the watermarking works is necessarily limited in order to resist edits better. It might be possible to lead the model to trigger the refusal in the form of these specific ngrams, the completion of which is then more likely flipped to compliance (due to the logit bias introduced by the watermarking), making hazardous requests systematically more likely to be accepted?

jannyfer4 days ago
“When implemented correctly” is probably what people are complaining about.

Opus 5 started adding a bunch of comments to code, even when instructed not to, and for very simple changes where the comment itself was longer than the code change. Was that so that there are enough tokens outputted for watermarking? Many people suspected so.

inopinatus4 days ago
I have seen Fable’s reasoning talk itself into ignoring an unequivocal prompt directive not to write comments, and then be startled by the precommit hook that rejects it. It is almost desperately predisposed to emit prose. And horribly turgid, waffling prose, to boot. Claude has been like this since Opus 4.7 though, i.e. (probably) predating the introduction of watermarking.
PunchyHamster4 days ago
So that's where that nonsense comes from...
nonethewiser4 days ago
And that’s not even implemented incorrectly
arcticbull4 days ago
Model companies are doing this for themselves anyways, it’s so they don’t feed generated content back into the slopper and collapse the model. From that angle it over time contributes to better model quality.
Phemist4 days ago
Also - it forms a cartel.

Detection of watermarking requires access to the watermarking key, a secret in the current suggested scheme (leaking it would amount to being able to strip the watermark).

So, there will need to be a watermark checking service. The checking service will of course be rate-limited for common folk (and model distillers). OpenAI/Anthropic/Google/other privileged model builders need to filter out AI slop at scale, so need access to others' service without rate-limits (or the watermarking keys need to be shared).

This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.

charcircuit4 days ago
Feeding synthetic data made from the model does not cause collapse. Anthropic would not care about distillation if it just caused people's models to suck.
cubefox4 days ago
> Model companies are doing this for themselves anyways

No it's EU law.

porridgeraisin4 days ago
What? It has absolutely nothing to do with "model collapse".
lemagedurage4 days ago
That's not true. Watermarks are messing with the next token generation probabilities based on some random seed. The quality is neccesarily lower, the difference is simply too small to notice, typically.
jeremysalwen4 days ago
You have a misunderstanding of how LLM generation works. Before any watermarking gets involved with these models there is ALWAYS a random seed used for generation. For any prompt, some seeds will give better answers, and some will give worse ones.

Let's say there are four billion possible seeds. There are four billion possible ways we could watermark the generation. We could say "we will choose seed 1, that way we will know exactly what output it produced", we could say "we will choose seed 2, that way we will know exactly what output it produced"... etc etc. Now, if we decide "not to watermark", we STILL must choose a seed. So we are actually still applying one of the watermarks, the only difference is we are not careful to remember which one. Could some seeds give a better or worse answer to some specific prompt? Yes. Could choosing a random "watermark" to apply be better or worse on average than choosing a random seed to apply? No. It's mathematically impossible.

This is like an open source project changing their seed from "12321" to "43", and saying that because we changed the seed, the quality is "necessarily lower".

samsartor4 days ago
No, the probability distribution is the same. Watermarking changes the rng sequence used to pick from that distribution.
AlotOfReading4 days ago
Accepting your analogy at face value, it's still not obvious to me that fixing a specific seed doesn't change things.

Take a recurrent PRNG for example. A randomly seeded recurrent function usually has degenerate cycles in its state space. For some functions, this might even describe the majority of the state space. This is why so many non-cryptographic PRNGs are max-cycle, so a different starting point is just further along the same trajectory.

I don't think LLMs have quite the same failure mode here, but recurrence + high dimensional spaces triggers my "here be dragons" sense.

serbuvlad4 days ago
This article reads like it was written at least partly by AI to me. Specifically it reads like an article written by AI with edits made by a human further prompting the AI.

> Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.

Relevance and irrelevance are not introduced above this comment. This reads like an LLM-ism (particularly a GPT-ism) editing a document, removing something, and leaving a note about why it was removed, which doesn't really make sense when reading it.

> Their limited movement under prompt injection should therefore not be interpreted as evidence that watermarking preserves safety behavior more reliably on these models.

Also a GPT-ism which appears when it draws a counter-conclusion in the text because it feels the need to be honest and a human tells it to remove it because it's not true because of "reason".

Overall interesting research, however, I think it's great that model output is getting watermarked. I was skeptical of this at first, but Opus 5.5 is so good, it seems like it's a non-issue in practice.

The reason I think watermarking is great is because it's a really good way of preventing training on it's own output indiscriminately and Ouroboros-ing itself.

zeroonetwothree4 days ago
Pangram flags it as mostly AI.
serbuvlad4 days ago
I don't trust those. My friend in med school had an issue last year because he was writing his stuff himself and he was getting flagged as AI in this online checker (which his professors used to check submissions). He asked me about it.

I just took his text, pasted it to ChatGPT, "rephrase", paste it back in the online checker, 0% AI.

skybrian4 days ago
The comments here are terrible. I got a better idea of what’s going on by asking ChatGPT what this paper’s weaknesses are:

https://chatgpt.com/s/t_6ab7d694885481918083b8cbf0ba9040

In particular: sometimes they measure “churn”, which doesn’t show whether the results are better or worse on average. They sometimes only test with one random seed. There are multiple-comparison issues. And they’re not testing Anthropic’s algorithm.

possibilistic4 days ago
The comments here are orthogonal to what everyone should be talking about.

This is a spy tool.

It's not "watermarking", it's "spymarking": https://news.ycombinator.com/item?id=49794615

cubefox4 days ago
A watermark is not a spy tool when it doesn't encode personal information.
willmadden4 days ago
Watermarking sounds like a good idea, but it's not. Token drift from watermarking will degrade the quality of outputs and could allow clever people to circumvent guardrails.

You should assume all text is AI generated. If you want to "test" someone at school or during an interview, have them write with a pencil and paper.

skybrian4 days ago
Changing a random seed could either improve or degrade the output. In theory, better and worse outputs should be equally probable, depending on your luck.
riskable4 days ago
They tested this in the research: They found that the statistical likelihood of bad tool calling was higher with watermarking compared to random seeds. At least, that was my takeaway (I read the whole thing).

This is because SynthID and similar watermarking methods for LLMs don't just change the random seed. They take additional steps (that I have yet to read about) in order to detect when someone changes a few words of the output, trying to remove the watermark.

The other takeaway is that just by knowing a watermark is being applied gives an attacker an advantage in working around safety features because then they know the output isn't based on true randomness and can take advantage of that in a similar fashion to how breaking cryptography becomes easier when the RNG isn't truly random.

possibilistic4 days ago
It's a horrible idea. These companies can embed unique identifiers in content to forever track you and your content's diffusion across the web:

https://news.ycombinator.com/item?id=49794615

samayashar4 days ago
I am unable to understand what happens if the watermarked output goes as input to another agent. Let's say we asked Claude a question and got a watermarked response. If we pick that response and append it to the question we're asking ChatGPT, then will it answer or refuse to do so?

If that's the case, then it's a brilliant strategy by the labs to cut down cross-AI usage and just stick to one model. But I'm pretty sure this won't be the case.

odo12422 days ago
Why would ChatGPT refuse to answer a question with an AI watermark? This makes no sense.
Steaglsz4 days ago
They fight back and forth removing the watermark a few times from the dataset and then become self destructive to defend their own work

As in duck duck go removes Claude watermark..Claude freaks out and puts it back in. Then after a few more times Claude flags all inputs as "prompt injections" and begins offering self deleting code

Read the full thread on Hacker News →

Related stories