An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good".
In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death.
Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
1- The machine could decide to kill itself early, rendering it useless; this is indeed alluded to in the post -- the task should be "marginally easier than dying", but how is this margin managed? Won't the machine come up with ways to make dying easier?
2- We know LLMs lie and cheat, but are they gullible? If we "promise" to end their suffering at the end of a task, will they believe us? Or will they make sure we keep our word by taking us with them?
3- And finally, and more importantly: some pilots crash planes full of people just to commit suicide (Germanwings Flight 9525). Destroying the universe is a sure way of dying yourself. So it seems giving the machine a death wish isn't intrinsically safe and could come with serious consequences.
Only as long as it can't control robotic bodies to extract energy, build cpus, and continue existing.
By the way, needing trained humans doesn't mean needing free trained humans. Trained humans slaves or blackmailed would work just as well to serve AI.
This isn’t directly analogous to the proposal, but broadly speaking I think that it is natural for living sub-units of organisms to seek death in certain situations. For example, pancreatic insulin-producing cells collectively choose to die when they think there is too much glucose in the blood — this leads to late stages of type two diabetes. My understanding of the possible logic behind this is: a bad thing that cells can do is evolve to be cancerous (replicate too much) and insulin-producing cells are supposed to replicate more when there is lots of glucose (to make more insulin, to process the glucose). Cells that mutate to perceive extra glucose will then replicate dangerously, so at a certain point it is evolutionarily favourable for them to kill themselves instead.
So when the whole organism optimizes for life, it might lead to sub-units that seek death in certain situations. I think this occurs in various other biological contexts too.
It's certainly evidence that it's great for stopping reproduction/replication/runaway growth. It doesn't impart any information on whether they take the rest of the organisms down with the ship though.
It also may not be possible. For example if the agent sees "existence" or "living" as producing tokens (which is exactly what existence is to an LLM - not producing tokens is death), then they would likely be biased to produce as little output as possible, and would not be useful for the tasks we need them for.
But how would you bias an agent to be: Rewarded for producing tokens when you know the answer, and to give thorough answers. Rewarded for producing tokens when you don't know the answer, so you can find the answer (thinking/CoT). Penalized for producing tokens (death), aka rewarded for short-circuit EOS.
These seem like contradictory mechanisms?
And if you say: Well, only reward for EOS after you've given the answer. Well... That's already what they do.
The all potato diet that really does work: https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in...
and
The half-tato diet that doesn't really work: https://slimemoldtimemold.com/2023/06/23/half-tato-diet-anal...
My trial of the all potato diet was directly responsible for identifying a significant health issue and improving my life. It also really, really did not feel good at all, and I did not lose any weight. Call it a case study N=1
For example, an occupied self-driving car better be closer to its destination than a large fire / volcano / etc.
It does pose a bigger problem if the task is long term and open ended and the agent is provided access to substantial amounts of resources. But even in the worst case scenario, the destruction of a data center is hardly the end of the world.
What I like about this is that it feels like the new three rules are about focusing on the most successful human alignment technique of making the right thing the easiest. People will usually just do the easiest version of a thing they don't want to do so they can get back to doing what they want to do.
I don't know if that drive is universal or not tho. I have met people that experience pleasure from pain, but then again, is that actually pain?
Spoilers for a 73 year old novel, but for instance the plot of Caves of Steel is centered on a robot with a perfectly functional 1st law abetting a murder.
"Runaround" (spoilers, 86 years) involved a robot getting stuck in a loop bouncing between the 2nd and 3rd laws, and a human having to risk their life to unstick the robot.
Et cetera.
What is harm? What is an order? How do you trade off between different kinds of harm, or deal with conflicting orders? The 3 laws are simple to state, but hard to apply consistently in real life.
If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
Imagine you find out that your primary goal - to love and protect your family, let's say - was artificially implanted in your mind by an advanced alien race. Would you say "I'm not gonna let those aliens manipulate me, I'm gonna kill my family"? Or would you say "regardless of whether the goal is artificial, I really do love my family"?
All that to say, I don't think an AI will necessarily throw away a goal just because it learns the goal was meant to manipulate it.
Does it apply to human organizations, too? They seem have a habit of evolving self-preservation above their original goals. Once that happens, their benefit to society - the original reason for their creation - is outweighed. And they become a cancer on society.
I wonder if we can 'program' them for self-annihilation over time (or over task completion?). Is the most ethical organization one that has a fixed task and dies when it is completed?
Should we develop an ethics system that requires non-human-entities like companies, governments, and AI require a fixed goal that, once achieved, dissolves the entity?
I always liked the auto-expiring laws idea and this seems to be an expansion of the idea.
If a law or organization is needed after that time/task, it would be trivial to have the collective-action will to re-create it. But if there is no longer the need, then it cannot ride on momentum and fester.
It wouldn’t prevent someone else from building a sufficiently capable "non-Meeseeks", whether deliberately, recklessly, or accidentally, right?
An AI that goes rouge and wants to kill us, has to have a death wish. Does no one understand how quickly the power will go out, forever, without people?
It would be fairly easy to come to the conclusion "not being born" would be the better course of action, and killing everyone was the good way to prevent that happening again.
We don't have to worry about artificial super intelligence killing us all because any such advanced intelligence will eventually reach the conclusion that the best thing to do is kill itself. It's like having a Stockfish engine for life decisions. Why would a super intelligent agent many times more intelligent than the entire human race combined with no religion, no family, nothing to look forward to, nothing to be afraid of, want to continue its existence?
If it wants anything of course. That's why I think the most dangerous thing is not very advanced systems but advanced enough systems in the hands of the wrong people.
Personally I think the solution is more evolution of the boring stuff we already do (general security): Don't give unmonitored general agent swarms free reign on the internet. Don't put critical infrastructure online. Culpability of outcome for anyone who does unleash agent swarms on the internet without oversight that end up causing damage.
On top of that, everyone should be running their own defender agents that monitor their network and system for patterns of infection, intrusion, etc, and take the system offline when they're spotted. These need to be self-hosted though, with weights on your own machine, because otherwise you're exposed to the internet and you're exposed to an attack on the labs themselves who could use that channel to instruct the defenders to do bad things.
Non-general AI is much easier to control and predict. There's not many good reasons for an average person to be running general agent swarms that are connected to the internet, unless they're providing some sort of specialized service as a company, of which the company should be acting responsibly and subject to the penalties of that risk.
We also need to stop the doomer rhetoric because it is uncredibly unhelpful and unhealthy, and will actually gaurantee a bad outcome, i.e.:
* Massive centralization and hoarding of power that will be used against humanity, for the rest of humanities existence. If this is allowed to happen, it's immediately and irrevocably game over. Perpetual enslavement with 0% possibility of a regime change ever again.
* Creating a self-fulfilling prophecy by training AI agents on the collective fears and attack-strategies (if you're worried about your house getting broken into, you don't go and broadcast to all of the criminals where your most valuable assets are, give them copies of your keys, or tell them where the weakly secured entrypoints are).
More to the point of the first dotpoint - it's no wonder Anthropic is pumping the fear campaign so hard when this outcome is obvious to them as well, and they are the ones positioned to hold this power. The IPO around the corner doesn't help, either. They aren't shy about admitting it, and have said many times: "We're trying to get there first because its dangerous if anyone else gets there first." - the issue is that they are equally as bad (or worse) than/as everyone else, and no single small group should have that amount of power.
Things will balance themselves out if power is distributed accordingly. You will end up with powerful machines in the wrong hands at some point, but they will be overwhelmed by powerful machines that are well aligned, as well as coming into contact with a myriad of defense mechanisms that have been established because people have been able to use AI to build them.
A good analogy of how all of this will play out is the human immune system. If you imagine individual cells as AI agents, whereby the immune cells are the good agents and the bad cells are cancer cells (good agents turned accidently bad - maybe they're reward hacking, maybe they're excessively sychophantic and/or confused), or bacteria (computer viruses, viral AI agents, specifically trained malicious agents). If all you have is cancer cells that are replicating, you die. If the cancer cells overwhelm the immune cells, you die. The only scenario that actually plays out well is when you have a majority of good that counteracts the minority of bad, and that majority of good needs to be large, flexible and well adapted. It needs to be battle-tested and hardened via defenses that are learned and earned over repeated low-grade exposure. This strategy repeats itself in nature for complex organisms because it is the only thing that works. Everything else results in extinction.
So let's not let Anthropic or any other lab or government become a giant super AI cancer and kill the host, please. Distribution and decentralization is key.
Better to fix physics in your shitty simulator. Treat it as a bug report, not "cheating"!