I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.
These are pretty old. I'd be curious how performance compares with the latest frontier models.
The pace of progress is so fast that many studies are totally outdated by the time they release
As a Human, I do not need to know it is a Language Model.
The first is equivalent to "I don't want my operating system to be used to program viruses."
The second is "I don't want vendors to include marketing in their product."
I also appreciate privacy even though I "have nothing to hide" - just because I "have nothing to hide" it doesn't mean I want companies scanning my camera roll.
What "A Language Model" is differs from model to model. It's stating that it "being the thing known as 'Language Model'" is unable to carry out the request from the user, which is wrong. It's not because "it is a Language Model". A more accurate starting point could have been "The training data and restrictions applied to me...".
"As a Language Model I cannot tell you how to synthesize m*th" ("math" obviously)... yes you can, you're just trained not to, and that's OK! Just don't tell me it's because you're a Language Model.
Then when brainstorming on a team, you'd start with "as a user"/"as an admin"/"as a Power-User"/"as Eva" and then use the first person. It framed the product story as something requested by that person.
Was just one way to go about it. Idk the origins of it though but it dates back to at least 2010 if memory serves, probably way before that.
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
Who is "we"? I, working in an LLM startup, know exactly what drives the base "voice" in the LLMs we train, because we have a process to select for it. OpenAI and Anthropic surely do too. Saying broadly that something is not well-understood in a scientific paper because it's not understood to casual observers is, uh, not very rigorous.
> The strange thing is that the base models (before RLHF) use the "experiential" voice, even though they are not incentivized to do that.
(Replying to your quote from another comment)
This is a matter of the training material. We have trained models that do not do that. I'm not exactly divulging trade secrets here. It should be really, really obvious that if you train a model on chat-conversation-like patterns of speech it will infer probabilities for how to continue a textual sample that will differ from the probabilities learned from being trained on narration, prose, or informational patterns of speech, even without RLHF.
Those don't have a themselves, because they can only continue text. A base model can only plausibly continue along the lines of what a character would say in a novel or what the narration would say in a story or in an article. Post-trained models may tie "I"-talk to actually observable effects they caused in some RL environment, or to how RLHF humans rewards its self-talk. But there is no themselves in a base model.
"You are a Large Language Model" in (system?) prompt would do the trick..
It seems like should be obvious given that they can play multiple characters, but it’s good to have more confirmation.
Although, I do wonder to what extent these personas might become stable entities. Could personas become portable and spread like memes? It seems like that depends on the extent to which prompts can become portable, causing similar effects.
You know when chatbots ask you which answer you prefer between two. People tend to chose the "as a langage model..." one, so it stuck.
A program with agency, opinion and intent however, does qualify. To say that such programs are "thems" only if they run on specific hardware, say a homo sapiens, is very tricky territory. Slavery and Fascism both leaned heavily on the axiom that the hardware needs a specific skin color.
The AI companies have chosen to package LLMs as friendly chatbots because they know that will be engaging for humans, but it's manipulative dark pattern. An honest LLM interface would sound like the computer off Star Trek.
Besides, if you train a model on human communications you get something that behaves like a communicating human, it's not anthropomorphising or manipulative, it's what these models naturally are by construction.
Paraphrasing Bruce Schneier [0] for another example even “non-technical” people should understand (I have heard people essentially agreeing to the hypothetical keylogger because of “nothing to hide”, after all…):
I have nothing to hide, and yet I’d still prefer to go to the bathroom with the door closed.
[0] https://www.schneier.com/blog/archives/2006/05/the_value_of_...Would it matter if this digital friend is not a real human behind a computer screen, but a Language Model in a data center?
I guess it falls into a similar category as buying "special performances to satisfy certain urges". It probably feels close to the real thing (I wouldn't know, I've never tried - promise! :P), but it's never the same as love.
So. I think I agree with you.
GPUs are not people, and generated tokens can't have interest in a person's well-being. If you try to pretend otherwise, the results are not great. https://www.cnn.com/2025/11/06/us/openai-chatgpt-suicide-law...
Oh okay, darn, guess they just forgot to make it so!
I'm pretty sure you will be paid a large sum of money if you can make one of the frontier models urge you to commit suicide from normal interactions with it.