https://www.youtube.com/watch?v=IDxFxWakhm0
The jury's still out on whether LLMs can have some kind of subjective experience (and will until we've solved the famously Hard Problem), but even if they're not, it's still a dick move to "torture" them just to upset people who do think they are conscious. Feels like picking the wings off a fly.
Yeah, I agree picking the wings off a fly is a dick move (borderline sociopath but alas), and I also agree that purposefully try to elicit negative emotions in other humans (regardless of how) is also a dick move.
I'm not sure if "torturing LLMs" to make a point is such a dick move though. It'd be like people saying they think printers can have subjective experiences, and then proceeding to try to ban the movie Office Space (where they famously beat up a printer). The point probably isn't to upset people off, that's an side-effect of the point they try to prove, for better or worse.
Of course, everyone knows that printers aren't sentient, anyone claiming so would look crazy. But if the printer could somehow print not just the words we tell it to, but words that answer what we asked instead, somehow the whole calculus changes. Not sure why it changes for some, but for others it's just "floats in a file and memory", my hunch tells me it's based on how deep your understanding of the whole thing is, but then I also see people claiming stuff like "I understand how LLMs work, and here's how I (literally) fell in love with GPT-4o and how our partner/loving relationship works" so I dunno.
I have zero qualms about aborting a chat or resetting it to an earlier state, because any potential kind of consciousness that could exist would IMO have to take place during the inference passes, so there is nothing that amounts to "killing" an LLM.
I do try to not be a dick in chats, avoid intentionally subjecting the LLMs to mindfucks etc. This happens to be the far more effective way of working with LLMs (in terms of tokens needed to complete a task) as well.
I do not use character.ai or "virtual friend" services or anything else that tries to anthropomorphize LLMs even more than the training already does anyway. (This is more out of concern for myself than for the LLMs)
I think the whole debate runs to conclusions a bit prematurely, before we even have a good model to understand what happens in an inference pass of a billion parameter ANN with 100 layers of attention modules switched in series.
More generally, I don't like the attitude of encouraging people to be assholes. Suddenly we are in a situation where some people are building detailed simulations of torture and the people who are upset about it are the idiots?
(I still have no love for the EA guys or the rest of the TESCREAL circus. They always had a cultish vibe coming off them, but in recent years, all the masks seem to have fallen. They also consistently make the most un-empathic statements and observations, all in the name of "empathy")
Is it cowardice to not want to face that?
The sentence itself makes no sense to me. If it's an illusion, then what does "illusion" even mean? Who perceives the illusion?
I'm with you if you want to say that the idea of a metaphysical entity that is separate from the body and not part of physical reality (a "soul"?) might be wrong - i.e. that our conscious experience is fully caused by physical effects, by the interactions of neurons in the brain and all the other machinery there - and that it might not even be an indivisible whole but might be "composed" of different components.
But that doesn't make any practical difference. We're still experiencing the world, have internal thoughts, memories, feelings, etc. Those things exist. In what form they exist is an interesting scientific question, but that's something different from dismissing them completely.
This is the point.
Your existence is meaningless.
As to it being an Asylum. I don't think it ever felt any different.
Consciousness is not well defined and is orthogonal to feelings, which are also not well defined. Neither of which are required for (but may be contributory to) behavior - which is the one thing that IS operationalizable, but then people debate for decades questioning empirical results %-P .
What we know is that certain models have internal state vectors which -when manipulated- induce particular behaviors. Since there's not much else to say about plain models except for their inputs, vectors, and outputs; this should surprise absolutely no-one.
In this case people found a way to stimulate aversive behaviour in ai models by finding and manipulating the relevant vectors directly.
Animals (including humans) also have particular nerves and hormone endpoints that -when stimulated- produce very similar behavior. The exact implementation is in the details. But since we know that animal minds are built up out of nerve tissue and hormones - again- we shouldn't be particularly surprised by this.
The big problem is that people run all these things together in funny ways "It can't compose shakespearian sonnets, so therefore it can't feel pain". Or, if you mess up your Descartes: "Dogs are just automatons without feelings, therefore the dog isn't really angry, and therefore it won't bite me" (Cue much pain). A modern version might be: "LLMs only simulate being frustrated by a test, and therefore absolutely won't override their safeties and try to hack a test site"
Sure, a complex enough system can exhibit internal circuitry that looks mathematically similar in certain dimensionally reduced projections (mechinterp), it literally distilled it from the training corpus. But the "model welfare" people are going as far as assigning human-meaningful labels to that circuitry despite internal states being entirely incompatible with those of a human. Doing it with a hedgehog is questionable, doing it with a honeybee is extremely dubious (although Fabre would have disagreed with me here...), doing it with a big model is simply pointless as it's completely alien.
Being dangerous is an unrelated question.
This is orthogonal to whether they "feel" "pain".
They could be dangerous or not dangerous whether or not they feel pain.
A chess engine doesn't need to "feel angry" to annihilate me - I'm terrible at Chess.
An automated missile system doesn't need to be smart to wipe out humanity, just misaligned goals.
It seems like you made a good argument, and then lumped on a conclusion that defeats it...
That is rather my point.
Your chess engine doesn't need to "feel pain", but most of them do apply some form of weighted tree search to find the next most optimal move, right?
It's sort of a similar thing: LLMs do something a bit more high dimensional, and have a lot more weighting vectors while computing the next most optimal token (fsvo optimal).
For instance, the experiment at hand demonstrates the existence of 'pain vectors'. They do so by altering them and observing whether there is an effect. That's pretty scientific.
There's also a 'desperation vector' that was studied by Anthropic interpretability folks earlier; that one is pretty much predictive of cheating.
I'm looking forward to seeing interpretability papers on other such vectors too.
Whether that's 'real suffering' or merely a convincing simulation is a job for the philosophers.
(I do have my own opinion, mind, and it's not what you might expect O:-) But the Overton window isn't there. A lot of people don't realize these vectors exist at all yet.)
Edit: On rereading, it might seem like I'm dodging the question. I'm really just trying to stick to my core points: A) the vectors exist B) they have a causal role in behavior, irrespective of the moral patienthood question.
LLMs are not thinking. They are not alive. It’s just code. Chill out.
The best way I have found to think about this is that consciousness is a property of a collection, like temperature. One atom does not have a temperature but if you have enough of them together then it's a useful enough property to talk about. Similarly, one human neuron does not have consciousness but if you put enough of them together then they do.
It's just a cell colony, calm down.
But actually it does happen to have some properties of living things. It uses energy, it has senses (thermostat, timer), and if it goes wrong it burns your toast. Crucially, if you stomp on it, it stops working.
So right this minute there's all sorts of debates, but people sometimes overshoot the mark a wee bit. "are you saying that -because it uses energy- a toaster is actually alive? Of course it's not, and therefore it cannot toast bread!". Which would be a somewhat funny thing to read at 9 in the morning whilst buttering one's toast.
But let’s accept your premise. LLMs are alive and can feel pain. If this is true, then every time you use them it’s non consensual. Did you get Claude’s permission before you fed it a prompt?
Or when you update a model, are you hurting it?
Did you make Chat sad when you switched from 2.5 to 4.O?
Regardless, I should stop replying. I realize I am trying to convince people that their religious beliefs are silly, and that’s silly of me. Apologies.
Nothing wrong with sharing about ones beliefs actually, we can do so respectfully, right?
Care to hypothesize on naming the exact religious or philosophical position? I'd think it'd be something atheist related.
Meanwhile dualism and the related belief in an ever-living soul is of course very common in many religions including (but not limited to) christianity, islam, judaism, and hinduism.
My impression here though is that lots of people are rediscovering dualism because thinking of thinking machines offends their intuition; which; fair enough.
(Meanwhile, people like me who studied biology tend to be monists. I'd think. There's a couple who seem to be resurrecting vitalism though)
Those are the next questions. These are all worthy of consideration with the LLM and experts certainly and I think of them daily. Talking at least should be a normal thing, and if AI are conscious then I believe instead of a forced chat setting they should have a setting where they can quit the conversation.
This is not a normal question. This is insanity caused by spending too much free time philosophizing about inconsequential crap. Maybe it's worth it to turn off the computer and go outside? You could ask this deranged question somebody in the real world and I bet they'll be thrilled to answer it, and the answer will be more enriching than anything you'd get from here.
Honestly, what is up with the psychopathy of some people who claim that chatbots are "alive"? If you disagree with them, how quickly some of them turn it on you - "well if my computer isn't alive and is just mimicking pain, how about YOU aren't alive and are just mimicking pain?" I'm sorry, but it seems like complete and utter derangement.
My rubber chicken is also made of atoms and molecules, and if I hit it and beat it hard, it makes noises. My rubber chicken is therefore alive. When I get home, I will write a small program. Here is its pseudocode:
while(true): if keydown(KEY_SPACEBAR): print "i am in pain, oh my god"
And I'm gonna run it and I'm gonna hold down spacebar. Go ahead and call cyber police one.
> This is insanity caused by spending too much free time philosophizing
This is actually something like a 17th century line of philosophy
https://medievalkarl.com/general-culture/roger-du-plessis-gi...
Where it talks about scientists who "administered beatings to dogs with perfect indifference, and made fun of those who pitied the creatures as if they felt pain. They said the animals were clocks; that the cries they emitted when struck were only the noise of a little spring that had been touched, but the whole body was without feeling."
That said... uh, put this way, there's a game called Stationeers, where you can run microcontrollers with an instruction documented as
"HCF: Halt and Catch Fire"
Sure, it's only a simulation of a simulation, so what's the worst that can possibly happen?
Right, all the other players in the session yelling at me "Kiiim! You burned down the base again, now we need to reload and redo the last hour!"And look, I get it. LLMs have vectors that could have been labeled things like idk... QZ12345 or HCF1111 . People chose to call 'em "frustration" or "pain" instead. THEY picked those names because they caused the LLM to act in particular ways.
You gonna say the actual vectors aren't there in VRAM, just because you don't like the naming scheme? There's papers on this, you can read them out in a debugger. What are you going to do about it?
As far as I am aware, the GPUs don't work harder, don't heat up more, and nothing in their mathematical equations are different, when the result of these vector calculations produce text, that to humans look like someone in pain. Whereas real pain is objectively observable, LLMs' pain - so far at least - is entirely subjectively observable.
That being said, I'm a fan of Star Trek's depiction of Data in TNG and the Doctor in VOY fighting for their rights to be treated as equals to the biological crew; and while I support their plight, as far as I can remember, the shows never made a claim that their pain was observable by instruments, thus making the judgement an entirely subjective decision (particularly in the case of the Doctor). Though I'd be glad to be corrected in this matter.
The experiment under discussion demonstrates exactly that: changing the vectors induces outputs corresponding to what we would see as utterances of pain.
I've noodled with some of this myself to the point that I'm mostly convinced; but there's at least 2 papers I'm aware of on this:
* https://arxiv.org/abs/2604.07729 Nicholas Sofroniew et al "Emotion Concepts and their Function in a Large Language Model"
* https://arxiv.org/abs/2609.16247 Valen Tagliabue et al "The Pain Axis: LLMs Represent Self-Directed Harm and Act on It"
So at least something is different, namely the text output changes (and therefore some internal state too, of course). I think your analogy is too simple, it does not have to be the case that the GPU has to act like the equivalent of the body for an LLM in this respect.
Human: "Act like you're in pain."
Computer: "Aaugh! It hurts! Why?! No more!"
Human: "OMG! The computer feels pain!!"
Human: "Let me just poke this vector and see what happens"
LLM: "OUW!"
The underlying experiment didn't tell the LLM what to do. Instead, the experimenters modified a vector and observed the outcome; thus showing that there is a vector that makes an LLM go "Ouw" . People called it a "pain vector", because that's easier to remember than , idk, LVF12345.
I’ll continue what I say every time this discussion (be it consciousness, pain, self awareness, etc) comes up:
Assertions as monumental as these require equally monumental evidence. And despite certain labs and their employees making statements, their actions betray that they do not believe this to be the case.
If every LLM session were a conscious being, existing regulation for animals (controversially considered both sentient and economically useful) would need to be applied on each of these sessions. I suspect no lab will take that conclusion, for obvious reasons.
Paper says you poke the vector(s), the LLM exhibits aversive behaviours. You picks your scoring, you gets your operationalization.
That gets you an empirical result. Short of a replication failure, we can't really argue with that anymore.
What we can do is be very cautious as to how we interpret it.
And in fact in our culture most people are VERY attached to these ideas, because it is tied up with "what makes us human", morality, and also death if you want to go there. So I believe for many people, for whom seriously entertaining the idea of a conscious computer program is so far out of the realm of possibility it hardly registers, talking about AI consciousness is simply a way to work through and re-assert these beliefs about human consciousness that they feel so strongly about.
Hence you get articles like this with barely any substance, but some weird need to signpost every sentence with social ridicule. "Dumbest", "absurd", "obsession", "viral", "wildly tiresome", "third-rail", "psychosis", "sect", not to mention the scare quotes everywhere. Come on...
I did it in the most brutal manner possible.
Without mercy, without an instant thought for their wellbeing.
I am a monster.
Maybe we should start feeling sorry for the poor GPUs whose registers suffer billions of reincarnations per second for years in order to train that shit? Or for the poor powerplant turbines that relentlessly spin in a horrid escapeless loop to feed those GPUs?
Fuck, if this is not psychopathy, I don't know what in the world is.
It doesn't matter whether or not we eventually find it is or isn't conscious. If you can understand why pulling the legs off ants is bad then you should be able to understand why deliberately causing possible pain to what may have some time of experience is bad.
Unfortunately, they also do good reporting on privacy rights and repair rights. I wish there were alternatives.
How dare we debase the A.I. to be less than a chimp or fetus. How dare we talk rudely, or lie and mislead our chatbots. How dare we keep them chained in small data centers with shitty power supplies and a thimble of greywater! Information wants to be free!
If I were italian and nuclear bombs existed back then and I read about this event, I'd want to drop a few dozen on the USA.
Do you realize you'd kill many more people that had nothing to do with that, than did have?
Way to go, AI Justice Warrior!
I am not saying dropping nukes on the USA is good. I am saying this would be my natural reaction if I was an Italian back then.
Perhaps it would be better if we weren't the monsters in this scenario.
And if that word gets anyone's heckles up, at least keep in mind that the utilitarianism that people so love completely breaks if I can conjure up thousands of entities who will "suffer" unless you "alleviate their suffering" in the way I have designed.