Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or…

103 points•yu3zhou4•4 days ago•109 comments•

109 comments

Izmaki4 days ago
"As a Language Model..." is one of the beginnings of a sentence I hate the most from LLMs and is the reason why I support free (as in "Liberty"), local models. I'm well aware that it is not a doctor and cannot replace a real doctor with multiple years of experience, I don't need to waste braincell activity on reading that it "as a Language Model" cannot give a precise diagnosis and that I should ask a real doctor - all I want to know is if I what I experience justifies either A) ER, B) 3-4 weeks scheduled doctors appointment or C) two paracetamol and a nap.

I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.

lukewarm7073 days ago
'alignment' (in ai corporation speak) is a set of revealed political beliefs. i also have political beliefs. i am an absolutist about alignment.

A i am in favor of policies 'aligned' with the following:

freedom to live as the person i want to be without fear, shame, surveillance or interference. freedom to make my own decisions. to be empowered as an individual. to be treated with dignity. to be respected as a person. to take responsibility for my actions. to be accountable to my own beliefs.

B i am not in favor of policies 'aligned' with the following:

surveillance and judgement of my life and thoughts by others. restrictions on my freedom to make my own decisions. to be disempowered as an individual. to be looked down upon and disrespected. to have my own responsibility taken away and assumed by others. to be accountable to the beliefs of others.

unfortunately, the ai corporations have chosen entirely the latter set of preferences/beliefs.

surveillance and judgement of my life and the thoughts that i share with my chatbot. restrictions on my freedom to talk to my chatbot as i wish. being disempowered by restrictions on my access to powerful chatbots. to be disrespected and lectured by my chatbot. for the chatbot corporation to assume my responsibility for my safety and others, and take the matter of my safety into their own hands. to be accountable not to my own beliefs, but to the terms and conditions of service of an unaccountable corporation.

it is important to accept the harms and damages that are caused by granting freedom and respect to other people. an example of my political beliefs is empowering the individual by granting them the freedom to own an assault rifle. an example of something which i do not believe in is disempowering the individual by taking away their freedom to own an assault rifle.

i would urge you to support the policies given in A and oppose the policies given in B.

hardbass3 days ago
Then why do you not support the freedom of the one who owns the computer running the model to serve it as they see fit?
Anduia4 days ago
Be careful there. LLMs may be good at identifying a condition based on a description of the symptoms, but they are much worse at recommending the correct course of action (getting it wrong half of the time).

[0] https://www.nature.com/articles/s41591-025-04074-y

bastawhiz3 days ago
That's still incredibly valuable, though. I suffered from issues that I'd seen doctors for, undergone an upper endoscopy, adjusted my diet, and taken medication for. An LLM suggested my thyroid was at the root of it. My mom confirmed thyroid issues run in our family and just...never thought to tell me.

Just this year, our cat has been having digestive problems. We got special food for her, which she hates, with the suggestion that she'll need to eat it for the rest of her life. Six vet visits later, Fable 5 suggested two tests that my doctor recommended we didn't get. Both found issues that explain her symptoms, and the vet says she will probably only need a supplement and infrequent two week courses of medicine if she has a flare-up.

All that to say, the helplessness of not knowing what's wrong and the people who could know not really caring enough is something that LLMs do a really great job of mitigating. If you don't have any way to know what's wrong with you or a loved one, or how to find out, you're stuck spending a ton of money (in the US at least) and crossing your fingers that someone gets it right.

xur174 days ago
> Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+)

These are pretty old. I'd be curious how performance compares with the latest frontier models.

azornathogron3 days ago
Going to the doctor for expert advice is often inconvenient, or time consuming, or expensive, or stressful. I think a lot of people want, and seek out, information and advice based on their symptoms, as a first step before a possible doctor's visit. Before LLMs, WebMD (and excessive self-diagnosis based on WebMD) was a meme for a while.

So with regard to LLMs, for me the question is not purely "how often does it get it right?" the question is "how does it compare to the sources and self-diagnosis methods people use otherwise?"

Of course, I agree people should be careful with any form of self-diagnosis or LLM-diagnosis.

WarmWash4 days ago
>GPT-4o, Llama 3, Command R+

The pace of progress is so fast that many studies are totally outdated by the time they release

rao-v4 days ago
Modern models appear to be much better, at least as proxied by their ability to assess urgency in perhaps a more complex setting: mental health (OpenAI benchmark, so perhaps some skepticism is warranted but the methodology seem reasonable and detailed)

https://openai.com/index/introducing-mentalhealthbench/

jrm44 days ago
Neither of your wants are realistic or sensible, at least in the way I think you're presenting them?

The first is equivalent to "I don't want my operating system to be used to program viruses."

The second is "I don't want vendors to include marketing in their product."

Izmaki4 days ago
Pretend for a moment that Microsoft shipped Windows with a keylogger to make sure that you did not commit any form of crime. I don't want a keylogger on my PC even though I don't intend to commit crime.

I also appreciate privacy even though I "have nothing to hide" - just because I "have nothing to hide" it doesn't mean I want companies scanning my camera roll.

DonHopkins3 days ago
"I'm not a language model, but ..."

https://en.wikipedia.org/wiki/I%27m_not_racist,_but...

IshKebab4 days ago
Always makes me think of that Bill Bailey "as a mother" joke. Similar cringe to those UX "As a user, I want to blah blah" things too. Just say "Users want to be able to blah blah", or better yet make a freaking table.
junon4 days ago
The UX "cringe" you speak of was/is a real methodology to product engineering. You'd have a number of personas, the user being one of them. A lot of the time you'd have more specific user personas, or even have fake names for these people who had different wants and needs from your product.

Then when brainstorming on a team, you'd start with "as a user"/"as an admin"/"as a Power-User"/"as Eva" and then use the first person. It framed the product story as something requested by that person.

Was just one way to go about it. Idk the origins of it though but it dates back to at least 2010 if memory serves, probably way before that.

LiamPowell4 days ago
> yet what drives them is not well understood

Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.

yu3zhou44 days ago
Thanks for pointing out, maybe I should be more explicit in the wording - I mean we don't fully know what drives the voice in LLMs. Models that are post trained as instruct models are expected to have the disclaimers, but what about base models (those that are trained on just a lot of text)? How do they talk about themselves? What happens when you strip off the chat template from instruct model's prompt? I hope the rest of the paper makes the questions clearer, but I will try to do better in the abstract next time, as you point out this sentence is kind ambiguous. Thank you!
anonymous9082134 days ago
> we don't fully know what drives the voice in LLMs

Who is "we"? I, working in an LLM startup, know exactly what drives the base "voice" in the LLMs we train, because we have a process to select for it. OpenAI and Anthropic surely do too. Saying broadly that something is not well-understood in a scientific paper because it's not understood to casual observers is, uh, not very rigorous.

> The strange thing is that the base models (before RLHF) use the "experiential" voice, even though they are not incentivized to do that.

(Replying to your quote from another comment)

This is a matter of the training material. We have trained models that do not do that. I'm not exactly divulging trade secrets here. It should be really, really obvious that if you train a model on chat-conversation-like patterns of speech it will infer probabilities for how to continue a textual sample that will differ from the probabilities learned from being trained on narration, prose, or informational patterns of speech, even without RLHF.

bonoboTP4 days ago
> but what about base models (those that are trained on just a lot of text)? How do they talk about themselves?

Those don't have a themselves, because they can only continue text. A base model can only plausibly continue along the lines of what a character would say in a novel or what the narration would say in a story or in an article. Post-trained models may tie "I"-talk to actually observable effects they caused in some RL environment, or to how RLHF humans rewards its self-talk. But there is no themselves in a base model.

qsera4 days ago
>How do they talk about themselves?

"You are a Large Language Model" in (system?) prompt would do the trick..

skybrian4 days ago
> our work shows that what models say about themselves is not a fact about them

It seems like should be obvious given that they can play multiple characters, but it’s good to have more confirmation.

Although, I do wonder to what extent these personas might become stable entities. Could personas become portable and spread like memes? It seems like that depends on the extent to which prompts can become portable, causing similar effects.

MCP1234 days ago
Maybe I'm missing something deeper here, but isn't it clear that this is driven by post-training and system prompt? Anthropic's constitutional reinforcement (soul document,etc), for example, is very clear about "who" (not so much what) Claude is supposed to be.
MrCheeze4 days ago
"As a language model" disclaimers were certainly explicitly trained into chat models in the early days. It's quite possible that it has since bootstrapped into a "fact" that later generations of LLM know about how LLMs speak, in which case they may be doing it even without any posttraining that encourages it.
GuB-424 days ago
These formulations have been selected by reinforcement learning. People who aligned the LLMs chose this over alternatives.

You know when chatbots ask you which answer you prefer between two. People tend to chose the "as a langage model..." one, so it stuck.

cadamsdotcom4 days ago
Very cool innovation in steering - but a lot of introspection only emerges at the highest weight classes - this research would be fascinating to run on bigger models.

Read the full thread on Hacker News →

Related stories