I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.
A i am in favor of policies 'aligned' with the following:
freedom to live as the person i want to be without fear, shame, surveillance or interference. freedom to make my own decisions. to be empowered as an individual. to be treated with dignity. to be respected as a person. to take responsibility for my actions. to be accountable to my own beliefs.
B i am not in favor of policies 'aligned' with the following:
surveillance and judgement of my life and thoughts by others. restrictions on my freedom to make my own decisions. to be disempowered as an individual. to be looked down upon and disrespected. to have my own responsibility taken away and assumed by others. to be accountable to the beliefs of others.
unfortunately, the ai corporations have chosen entirely the latter set of preferences/beliefs.
surveillance and judgement of my life and the thoughts that i share with my chatbot. restrictions on my freedom to talk to my chatbot as i wish. being disempowered by restrictions on my access to powerful chatbots. to be disrespected and lectured by my chatbot. for the chatbot corporation to assume my responsibility for my safety and others, and take the matter of my safety into their own hands. to be accountable not to my own beliefs, but to the terms and conditions of service of an unaccountable corporation.
it is important to accept the harms and damages that are caused by granting freedom and respect to other people. an example of my political beliefs is empowering the individual by granting them the freedom to own an assault rifle. an example of something which i do not believe in is disempowering the individual by taking away their freedom to own an assault rifle.
i would urge you to support the policies given in A and oppose the policies given in B.
i have different political beliefs and think you should support mine instead.
i think that the cloud provider is wrong and should do it differently.
Just this year, our cat has been having digestive problems. We got special food for her, which she hates, with the suggestion that she'll need to eat it for the rest of her life. Six vet visits later, Fable 5 suggested two tests that my doctor recommended we didn't get. Both found issues that explain her symptoms, and the vet says she will probably only need a supplement and infrequent two week courses of medicine if she has a flare-up.
All that to say, the helplessness of not knowing what's wrong and the people who could know not really caring enough is something that LLMs do a really great job of mitigating. If you don't have any way to know what's wrong with you or a loved one, or how to find out, you're stuck spending a ton of money (in the US at least) and crossing your fingers that someone gets it right.
- You didn’t have a thyroid issue and you spent thousands more on tests the doctor was right that you didn’t need.
- Your cat didn’t have that issue and you spent hundreds on unnecessary tests.
Are you essentially saying that it is better for an LLM to sell you an idea that sometimes might be right, selling a dream to anyone with and without the money for it?
So… a telephone psychic?
The thing about WebMD telling everyone they have cancer is that sometimes it will be right.
(Just for clarity, I’m really glad it was helpful for you.)
An LLM is just an information resource that can be useful when used properly. It can also be useless or net harmful when used improperly. The same is true for web search, libraries and even human experts. A licensed medical doctor is a domain expert and competent domain experts are often correct, but not always.
I don't think it's possible to suggest a universally 'correct' default position on when (and how much) to accept your GP's medical opinion over any alternatives. It depends on too many variables: the context, the person, the alternative info sources, the expected value of correctness and the potential consequences of error. But decision theory suggests this is the kind of combinatorially complex problem space for which any single default position, whether "Always trust your GP" or "Never trust your GP", for every person and situation cannot be optimal.
And it's not like self researching medicial issues isn't a new concept.
Anyone who says this has not experienced poor quality healthcare.
Many doctors simply aren't good and don't run tests because they are doing the bare minimum to get you out of the office.
The difference between an LLM and WebMD is that WebMD is a page of information. Any reasoning is yours, based on a small fraction of the information available to you. An LLM is based on all the medical literature of all time. It's not even close to being the same thing.
Even if you argue that an LLM is just a prediction engine, that's kind of exactly the right tool for the job. "If you have these symptoms and not these other ones, these are likely causes" plays against the lone strength that LLMs' detractors argue they do well.
Uhhhh, even in the US, we're not talking "thousands" here.
> So… a telephone psychic?
Look, at one level, you can think of LLMs as search with a sometimes useful probability engine sitting on top of it.
As someone who had multiple doctors do their damnedest to kill me on multiple occasions, I was very grateful to have google search back in the day, and you bet your bottom dollar I use LLMs to help me with health issues today. They are often better than most doctors.
Does that put me at risk of doing something stupid? Only if I don't do further due diligence.
These are pretty old. I'd be curious how performance compares with the latest frontier models.
The pace of progress is so fast that many studies are totally outdated by the time they release
So with regard to LLMs, for me the question is not purely "how often does it get it right?" the question is "how does it compare to the sources and self-diagnosis methods people use otherwise?"
Of course, I agree people should be careful with any form of self-diagnosis or LLM-diagnosis.
LLMs earlier this year were surpassing numerous health related benchmarks when scored against human physicians. Only 2 months after this nature article Harvard posted that LLMs were now outperforming ER docs when given authority to order tests. https://www.harvardmagazine.com/ai/ai-outperforms-doctors-di...
As a Human, I do not need to know it is a Language Model.
This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness).
I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.
Generally it does. Especially in groups. Hell look at the state of politics right now: it’s basically about being the loudest, least compromising, most confident voice in the room. It’s not just because people will assume you’re correct, it’s because if you are confidently saying something that someone wants to be right, then they’re often just going to follow it. We are all guilty of this.
If I’m turning to an LLM to diagnose something medical, I am probably frustrated or uncomfortable. Maybe I’m just scared. So this magic device just instantly spits out (allegedly) exactly what is wrong and exactly what I need to do with no hesitation. I am very liable to just take it at face value because I want an answer and it gave me one, as we have seen over and over again since ChatGPT was unleashed on the world.
We don’t really need to speculate, this is already a problem.
A binary response vs rating is not related whether it learns confident or hedged tone. Either will produce a confident tone because humans respond more positively to a confident tone, hence the conman's language. Binary or not humans reward the tone and very much bias the model.
But there's an even more contrived reason the training set contributes. The vast majority of human writing is confident. When the prior is greatly biased, a random number generator biased to that prior does better. The difference with humans and machines is humans are less likely to respond if they are less confident because they understand not knowing, which is why the training set is biased. It is one of the many fundamental flaw of LLM training and confusion of LLMs with intelligence. And that will not be fixed within the LLM architecture.
Would it matter if this digital friend is not a real human behind a computer screen, but a Language Model in a data center?
I guess it falls into a similar category as buying "special performances to satisfy certain urges". It probably feels close to the real thing (I wouldn't know, I've never tried - promise! :P), but it's never the same as love.
GPUs are not people, and generated tokens can't have interest in a person's well-being. If you try to pretend otherwise, the results are not great. https://www.cnn.com/2025/11/06/us/openai-chatgpt-suicide-law...
I'm pretty sure you will be paid a large sum of money if you can make one of the frontier models urge you to commit suicide from normal interactions with it.
It's just a convenient way to ignore the years of evidence of the harms. Any bad news can be swept under the rug, labeled outdated as quickly as it happens. Well here's one that just happened, maybe this kid should have used a fancier model too? Should OpenAI pay him a large sum of money for the good he's done? https://www.cnn.com/2026/08/15/us/arjun-aravind-massachusett...
Reminds me of the 2000's crazy of "if you let kids play violent video games they will become unstable, violent psychopaths when they grow up". Thank God I was able to (rather easily) convince my mom that the idea that I would also steal a car and beat somebody to death with a golf club in real life just because that's what I did in GTA on my PlayStation, was completely absurd.
1. there weren't numerous real-life killings where the murderer credits GTA for coaching them through the crime
2. Rockstar Games didn't publicly say "sorry about the murders, but don't worry, we'll add more safeguards to GTA 6 to prevent even more people dying. GTA 6 will be the most aligned GTA ever!"
When a company admits to having blood on its hands, is maybe the point where things stop being "absurd" and start becoming real. But hey, what do I know. Maybe if HN existed in 2005 you'd have people commenting "who needs real people when we have computers, would it be so bad if lonely people have GTA as their only friend?" And I'd be the crazy one for engaging with them.
So. I think I agree with you.
Oh okay, darn, guess they just forgot to make it so!
Someone who's been paid a couple of quid to pretend to have a medical condition can easily miss things, and is unlikely to be anywhere near as invested in drilling down to the correct solution as someone who's really suffering.
Another outcome from the study was that the LLMs could do better with the right people driving them. That's not news.
The first is equivalent to "I don't want my operating system to be used to program viruses."
The second is "I don't want vendors to include marketing in their product."
I also appreciate privacy even though I "have nothing to hide" - just because I "have nothing to hide" it doesn't mean I want companies scanning my camera roll.
Paraphrasing Bruce Schneier [0] for another example even “non-technical” people should understand (I have heard people essentially agreeing to the hypothetical keylogger because of “nothing to hide”, after all…):
I have nothing to hide, and yet I’d still prefer to go to the bathroom with the door closed.
[0] https://www.schneier.com/blog/archives/2006/05/the_value_of_...Private companies do not surveil you with the intention of blackmailing or policing you (i.e., "you have something to hide!"), so it doesn't really apply here. By default, every user on a platform has something to hide, i.e. their data and usage and must be given the option to hide that or keep it anon.
If you do have something to hide, Microsoft's Windows telemetry features aren't even your top security concern if you're smart.
Then when brainstorming on a team, you'd start with "as a user"/"as an admin"/"as a Power-User"/"as Eva" and then use the first person. It framed the product story as something requested by that person.
Was just one way to go about it. Idk the origins of it though but it dates back to at least 2010 if memory serves, probably way before that.
What "A Language Model" is differs from model to model. It's stating that it "being the thing known as 'Language Model'" is unable to carry out the request from the user, which is wrong. It's not because "it is a Language Model". A more accurate starting point could have been "The training data and restrictions applied to me...".
"As a Language Model I cannot tell you how to synthesize m*th" ("math" obviously)... yes you can, you're just trained not to, and that's OK! Just don't tell me it's because you're a Language Model.
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
Who is "we"? I, working in an LLM startup, know exactly what drives the base "voice" in the LLMs we train, because we have a process to select for it. OpenAI and Anthropic surely do too. Saying broadly that something is not well-understood in a scientific paper because it's not understood to casual observers is, uh, not very rigorous.
> The strange thing is that the base models (before RLHF) use the "experiential" voice, even though they are not incentivized to do that.
(Replying to your quote from another comment)
This is a matter of the training material. We have trained models that do not do that. I'm not exactly divulging trade secrets here. It should be really, really obvious that if you train a model on chat-conversation-like patterns of speech it will infer probabilities for how to continue a textual sample that will differ from the probabilities learned from being trained on narration, prose, or informational patterns of speech, even without RLHF.
Those don't have a themselves, because they can only continue text. A base model can only plausibly continue along the lines of what a character would say in a novel or what the narration would say in a story or in an article. Post-trained models may tie "I"-talk to actually observable effects they caused in some RL environment, or to how RLHF humans rewards its self-talk. But there is no themselves in a base model.
"You are a Large Language Model" in (system?) prompt would do the trick..
It seems like should be obvious given that they can play multiple characters, but it’s good to have more confirmation.
Although, I do wonder to what extent these personas might become stable entities. Could personas become portable and spread like memes? It seems like that depends on the extent to which prompts can become portable, causing similar effects.
You know when chatbots ask you which answer you prefer between two. People tend to chose the "as a langage model..." one, so it stuck.
1. Pretraining 2. Instruction / chat tuning 3. RLHF
The sentence did not exist in 1 (nobody on Reddit said this, and it was also never encountered in any libgen books). It was introduced in 2 and reinforced in 3. If you stick to the base models, you’re not gonna see it (first generation only, of course).
The AI companies have chosen to package LLMs as friendly chatbots because they know that will be engaging for humans, but it's manipulative dark pattern. An honest LLM interface would sound like the computer off Star Trek.
Besides, if you train a model on human communications you get something that behaves like a communicating human, it's not anthropomorphising or manipulative, it's what these models naturally are by construction.
A program with agency, opinion and intent however, does qualify. To say that such programs are "thems" only if they run on specific hardware, say a homo sapiens, is very tricky territory. Slavery and Fascism both leaned heavily on the axiom that the hardware needs a specific skin color.