Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or…
109 comments
I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.
A i am in favor of policies 'aligned' with the following:
freedom to live as the person i want to be without fear, shame, surveillance or interference. freedom to make my own decisions. to be empowered as an individual. to be treated with dignity. to be respected as a person. to take responsibility for my actions. to be accountable to my own beliefs.
B i am not in favor of policies 'aligned' with the following:
surveillance and judgement of my life and thoughts by others. restrictions on my freedom to make my own decisions. to be disempowered as an individual. to be looked down upon and disrespected. to have my own responsibility taken away and assumed by others. to be accountable to the beliefs of others.
unfortunately, the ai corporations have chosen entirely the latter set of preferences/beliefs.
surveillance and judgement of my life and the thoughts that i share with my chatbot. restrictions on my freedom to talk to my chatbot as i wish. being disempowered by restrictions on my access to powerful chatbots. to be disrespected and lectured by my chatbot. for the chatbot corporation to assume my responsibility for my safety and others, and take the matter of my safety into their own hands. to be accountable not to my own beliefs, but to the terms and conditions of service of an unaccountable corporation.
it is important to accept the harms and damages that are caused by granting freedom and respect to other people. an example of my political beliefs is empowering the individual by granting them the freedom to own an assault rifle. an example of something which i do not believe in is disempowering the individual by taking away their freedom to own an assault rifle.
i would urge you to support the policies given in A and oppose the policies given in B.
Just this year, our cat has been having digestive problems. We got special food for her, which she hates, with the suggestion that she'll need to eat it for the rest of her life. Six vet visits later, Fable 5 suggested two tests that my doctor recommended we didn't get. Both found issues that explain her symptoms, and the vet says she will probably only need a supplement and infrequent two week courses of medicine if she has a flare-up.
All that to say, the helplessness of not knowing what's wrong and the people who could know not really caring enough is something that LLMs do a really great job of mitigating. If you don't have any way to know what's wrong with you or a loved one, or how to find out, you're stuck spending a ton of money (in the US at least) and crossing your fingers that someone gets it right.
These are pretty old. I'd be curious how performance compares with the latest frontier models.
So with regard to LLMs, for me the question is not purely "how often does it get it right?" the question is "how does it compare to the sources and self-diagnosis methods people use otherwise?"
Of course, I agree people should be careful with any form of self-diagnosis or LLM-diagnosis.
The pace of progress is so fast that many studies are totally outdated by the time they release
The first is equivalent to "I don't want my operating system to be used to program viruses."
The second is "I don't want vendors to include marketing in their product."
I also appreciate privacy even though I "have nothing to hide" - just because I "have nothing to hide" it doesn't mean I want companies scanning my camera roll.
Then when brainstorming on a team, you'd start with "as a user"/"as an admin"/"as a Power-User"/"as Eva" and then use the first person. It framed the product story as something requested by that person.
Was just one way to go about it. Idk the origins of it though but it dates back to at least 2010 if memory serves, probably way before that.
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
Who is "we"? I, working in an LLM startup, know exactly what drives the base "voice" in the LLMs we train, because we have a process to select for it. OpenAI and Anthropic surely do too. Saying broadly that something is not well-understood in a scientific paper because it's not understood to casual observers is, uh, not very rigorous.
> The strange thing is that the base models (before RLHF) use the "experiential" voice, even though they are not incentivized to do that.
(Replying to your quote from another comment)
This is a matter of the training material. We have trained models that do not do that. I'm not exactly divulging trade secrets here. It should be really, really obvious that if you train a model on chat-conversation-like patterns of speech it will infer probabilities for how to continue a textual sample that will differ from the probabilities learned from being trained on narration, prose, or informational patterns of speech, even without RLHF.
Those don't have a themselves, because they can only continue text. A base model can only plausibly continue along the lines of what a character would say in a novel or what the narration would say in a story or in an article. Post-trained models may tie "I"-talk to actually observable effects they caused in some RL environment, or to how RLHF humans rewards its self-talk. But there is no themselves in a base model.
"You are a Large Language Model" in (system?) prompt would do the trick..
It seems like should be obvious given that they can play multiple characters, but it’s good to have more confirmation.
Although, I do wonder to what extent these personas might become stable entities. Could personas become portable and spread like memes? It seems like that depends on the extent to which prompts can become portable, causing similar effects.
You know when chatbots ask you which answer you prefer between two. People tend to chose the "as a langage model..." one, so it stuck.
Read the full thread on Hacker News →
Related stories
- Hacker News · 2 points · 9 days ago
- Hacker News · 2 points · 1 day ago
- Mumble – Open-Source, Self-Hosted Voice Chatdigitalescapetools.comHacker News · 1 points · 10 days ago
- Hacker News · 1 points · 2 days ago
- Hacker News · 1 points · 9 days ago
- Hacker News · 1 points · 5 days ago