We can (and should) try to improve AI safety and alignment, but perfect safety is probably impossible without losing the capabilities that make AI unique. Black-and-white positions that insist anyone creating or using an AI is responsible for anything it does wrong under any circumstances are also unrealistic. AI is a general purpose technology, and we have always accepted that it is not practical for suppliers to envisage and constrain every downstream use: Car manufacturers cannot prevent every instance of mechanical failure, nor can they stop cars being used in crime.
This is because they can't take the blame, they can't be held responsible.
Decision laundering is a great term. There is ultimately a human who gave the program free rein. You can't say "sorry my car hit you, it's not my fault" especially when the person driving works for the car manufacturer too.
The question isn’t who or what made the decision, per se. It’s who is responsible for the outcome. The entity making the decision might be the actual cause but not the proximate cause (the cause most closely connected to legal liability) for the harm that occurred.
Related: The Unaccountability Machine (good book by Dan Davies, who also wrote Lying for Money about financial fraud).
This. I agree 100%.
Computers make decisions all the time, and have been doing so for decades. But we still keep the human behind the computer accountable.
Don’t confuse lack of accountability with lack of responsibility.
If one needs to be "allowed" to make a decision, does that even qualify as a decision?
For example, I can decide which cereal to purchase at the supermarket. But I cannot decide which cereals are offered to me at the supermarket. The supermarket has allowed/enabled to make a constrained decision.
So does it follow then that the agents in question do have "some" free* will? I'm not sure you realize what conclusion naturally follows from this, and how it contradicts the point made in the article and the top-level comment I replied to.
(*) My earlier question wasn't explicitly about free will but we'll go with it.
While I agree that computers cannot currently be held accountable for decisions (because I do not think they have free will), I think it’s transparently clear that computers make decisions all the time, and have done since they were first implemented.
Many of these incidents relate to hacking, where criminal law commonly requires intent.
Claiming that a machine cannot make a decision does not mean an employee at OpenAI decided to hack an Australian government website.
The AI labs say "AI can make mistakes". That's why the responsibility of using it is on you. This may change if we get more legislation. But given the argument above, it will probably never shift to the AI itself having any responsibility. Instead the AI labs can be responsible.
It's like having an assistant that can be helpful, but not always. We can bicker about the ratio of helpful to unhelpful interactions, but there will be some unhelpful moments.
And hoo boy, those can range from "eh, whatever" to being like a deranged toddler just got its hands on whatever tools and permissions were granted to the "mostly helpful" assistant.
And we can't get rid of that chance. We can pare it down a bit. But we can't get rid of it because it is fundamental to how the technology works.
Instead of "Anthropic submits false tip to police", I'd prefer "Anthropic employee John Smith deploys an agent that submits false tip to police". Of course, John Smith will be able to then say "I did so as ordered by my superior, Ms Sue", and so on, with a very clearly established chain of liability.
A corporation and an AI are not very different in many aspects, except one of them is very favourably treated by law, almost as a living person.
"Anthropic employee John Smith [killed the children with the gun||launched a cyber attack] because he was ordered to by his superior and [has all the blame because such orders are null || we agree that Anthropic can order children to be killed]"
A bad manager takes personal credit when things go well and finds a scapegoat when things go bad. A good manager gives credit when things go well and takes responsibility when things go bad.
It's wrong to scapegoat the agent. You, the human who executed the agent, made the decision to execute it. You gave it credentials, and set up insufficient guardrails. Don't scapegoat the agent for your poor decisions.
Likewise, it's good to have humility when you get great results out of agents. Yes, the agent made the good results happen. It's OK to tell people that your choice of agent is amazing and did an incredible job. In any half-decent work culture, that should rub off on you as well.
That's not the situation here. We aren't enabling the career progression of software which does matrix multiplications.
Accountability of us schmucks at the end of the org chart goes both ways. We get some reward when we do it right, and we improve when we do it wrong.
A reward might be something tangible or might be as small as your own personal satisfaction.
An improvement is hopefully identifying what went wrong and doing better next time, and in extreme cases your employer will let you try again at a different workplace.
The question of whether it's better to frame a reaction to issues in a positive, do-better-next-time or a negative, don't-do-that-again way (fwiw, I agree that the positive way is better) is orthogonal to the argument I was trying to make about who is the recipient of that reaction. Scapegoating the agent is farcical. Clearly, "make no mistakes" style prompts are a poor approach to agentic development.
> the manager's job is to enable career progression of their reports... we aren't enabling the career progression of [fancy calculators]
While I do agree that good human managers concern themselves with career development of their human direct reports, I disagree that it's a fundamental part of the job description. The foundation is for a manager to get and keep their reports being productive. Human workers are often concerned with career progression, so good managers concern themselves with such career progression as well. Agents are not, but that doesn't mean that lauding good agentic work is meaningless. Lauding the work is also valuable for: (a) reinforcing a culture of positivity and celebration, which rubs off on human workers as well, (b) part of a culture of continuous evaluation of agents and models - asking questions like, are there cheaper or faster models that would be just as effective?
Humans, especially ones under abusive managers, often have a lot of responsibility but no authority that affects the outcome.
In an ideal situation you either have both, or none. You make the decision and answer for it, or you follow decisions but don't get blamed when it's wrong.
LLMs managed to slide into a space where they get to make authoritative decisions but if something fails we don't really blame them.
I'm not saying LLMs should be held responsible, but it's funny how people claim they replace humans, but are giving them an extremely easy, kid-gloves, success target.
These firms have been using computers to make decisions about trading billions of dollars for over 30 years.
At the same time, I've been involved in hundreds of After Action Reviews of incidents and it's not always simple or easy to both figure out which party was ultimately responsible and also how to allocate the "dollar error budget" to the parties overall.
(This could be a blog post in and of itself btw)
Ambrose Bierce, The Unabridged Devil's Dictionary (1906)
...like it already happens in europe, or at least in Italy.
This is basic law and risk management in society. Let's not pretend we have just invented corporations and don't know what to do with them yet.
Oh, wait, that was a core concept from "Mein Kampf": parliament are bad because nobody is responsable, hence nobody care about doing the right thing ! (probably the only concept from that book that I can agree to)
One does not, under any circumstances, "gotta hand it to Hitler".
Regardless, this is less about dictatorship and more about personal responsibility and accountability
A solution, perhaps, would be to ensure that such decisions are not anonymous and that, if the decisions had bad consequences, then the related people would be accountable
Of course, this would make the politic business less attractive. Would that be a bad idea ?
But I don't think the average person with their $20 subscription to whoever is the primary source of revenue. Enterprises fund most of it by buying API keys and making those fancy chatbot buttons on their sites that nobody clicks, as well as some internal business automations if I had to guess.
If it was just average people with $20 subscriptions funding all this (keep in mind, most people using AI services do not pay for it), by now, all these AI companies would have gone bankrupt.
These businesses lose unbelievably large amounts of investor money. Their AI-related revenues don't get even close to paying for their AI-related outlays.
End user (invididual, small business, large business) clients are still complicit not because they are funding it, but because they are signalling to investors that this is something people/businesses want.
Nobody gets off the hook here. It is every bit as much end user nihilism as it is AI company nihilism.
AI spending is a bit like Pokemon games - most of your players are free to play, only earning you traffic and maybe some more players after recommending it to their friends - that's the average AI user. Around 5-10% of players pay a bit to the game - maybe to unlock some extra characters - those are the $20-$200 AI subscription users, with these you are not going to fund the whole business but it's some side cash. Then you have the 1℅, which spend massive amounts of money onto the game, and those are the ones primarily funding it - that's the business spending and where most of the money actually comes from.
You can only blame the average person for using the cloud service and recommending it to others.
Erm, you appear to have missed the last 100+ years of updates to Western Capitalism.
People don't want it, you make the demand with hundreds of millions of spend on advertising (brainwashing), influencers, and social proof.
The capital holders don't sit back and see which company wins so capital can be directed to the best solutions, they buy up the competition, they pay corrupt politicians, they use current v market placement to embed their products name into billions of workplace computers and spread stories about how it's going to 'increase productivity but never at the cost of jobs', etc.
Yes, ultimately a dollar is a vote, but I don't think you can blame the people being manipulated "every bit as much" as those controlling the manipulation to feed their avarice and/or megalomania.
Not really.
> Yes, ultimately a dollar is a vote, but I don't think you can blame the people being manipulated "every bit as much" as those controlling the manipulation to feed their avarice and/or megalomania.
Oh but I do.
Collective social guilt is a thing. e.g. who is to blame for drug violence? Not only drug dealers.
The problem is entirely different. Sometimes they make decisions we don't like, but unlike with people we haven't found a way to punish them that satisfies our sense of justice or discourages other computers from making similar decisions.
It does not matter if the CEO is an AI or a human. In either case, of course they can make decisions -- and be held responsible. Not in any emotional sense (that we seem to want to bake into the concept for added confusion), just that if you are not satisfied with the result, you can replace the human or AI.
I am not saying that this is a fun vision of anything, but why are we making it more complicated than it is? The parts are all there and explained.
My understanding is that 'emotional' sense emerged organically in LLMs and was not intentionally baked in.
As silly as it may seem, the ways emotions come through in our words do seem to have an effect on agents. Likely because they were trained on human writing, most of which cannot be separated from the emotions of the writers any more than your emotions can be separated from your logic & reasoning abilities.
While it seems reasonable that you need humans to take accountability, it looks like a very shaky position to assert that computers can't make decisions. And going back to chess computers, we acknowledge that computers make better decisions than humans in that area.
Forbidding computers that can make better decisions than humans from making decisions to keep accountability just seems to be arguing to hire fall guys. You still want the better decision maker to make decisions, but you need someone there to take the fall if things go wrong.
Very weird incentives.
You might have an org policy that says that an llm can't make decisions because an llm can't be held accountable - but that's your choice.
You don't need to operate with a model where everything has a human accountable for it in the end.
If you don't like the decisions an llm makes, you can replace it with another or put a human in charge of that decision in the future instead.
I'm not saying it's a good idea, but you can technically do it.
The piece: "Anthropic COULD stop their models from doing what they are doing, and they are actively choosing not to"
As if _Anthropic_ could make a decision.
Companies cannot make decisions! All of these incidents are the fault of people. Companies are not deciding to do this, people are. There is little more to it than that.
They generate text which resembles reasoning and may include a tool call
The harness runs tool calls
The human made the decision to write the harness
The human is responsible
Also phrased as ‘to make an apple pie from scratch, one must first create the Universe’
There is no end of the world, nor autonomous agent decisions nor any fantasies they are building up to make people believe they have a pandora box in their hand.
What we instead have are lies crafted to lure non tech people and keep this running as long as they can.
Oh and let's also not forget that they are allowed to illegally download basically as much as they want from libgen/annas archive/torrents to train their models under the "fair use" umbrella, yet mere mortals can go to jail for way much less.
If you are driving a car, it's pretty clear where the accountability falls. You control the car. You have free will. If the car hits a tree, you hit a tree.
What happens when a car hits another car? Things get murkier. There are lawyers who have devoted their careers to this. Sometimes nobody winds up at fault. Unavoidable accidents happen.
Now let's ask what happens when a leader orders their military to do something. e.g. Say a President instructs his troops to do whatever they think is necessary to get the job done, and they go and Abu Ghraib a bunch of prisoners. Who is responsible? The leader probably is, ultimately, but there are a lot of human beings with free will in between the leader and the people wielding electrical cords and pliars. Typically, somebody fairly close to the bottom takes the blame, even if most people suspect it should fall higher. Some low rankers will get court marshalled but the leader will have the occasional shoe thrown at him. That's okay. He's good at ducking.
Say you're running a war against people you really just want gone. Your military has a tradition of conducting laboriously researched precision strikes to take out people they don't like. Your people are working hard, but they're only managing a couple strikes a week, and you have so many more bombs than that. You don't even pay for most of them! What if you could just ask a machine to find you the leader's underlings, and their underlings, and so on, right down to raw recruits? Nobody has to check on the machine. Why bother? Just believe the machine and drop the bombs. If anyone ever calls it a war crime, you can just blame the machine, right? Well, not so fast. Most people are smarter than that!
Despite all the sci-fi we consume that portrays computers as having agency, malevolent or not, most people instinctively understand that, in the real world, computers are just tools, more akin to cars than armies. LLM's are just software that run on those tools. Anyone using a LLM can simply choose not to, or they can choose to be very careful in how they use it. There isn't a complex web of free wills to untangle. There's the machine and the person who controls the machine.
We currently have an administration that is desperate to look the other way while U.S. AI companies do things that any reasonable person would hold them accountable for. "Industry will regulate itself!" The economy mostly sucks because of this same government's complete lack of economic sense or even basic self-control, but the AI sector is propping things up. They're not going to regulate the golden goose unless the goose does something so unbelievably stupid that there's no choice. Wouldn't want to pop the bubble!
Yes, we are headed for more stupidity. The "tequila and handguns" kind of automated stupidity that only a computer and a truly feckless operator can give us.
https://www.hec.edu/sites/default/files/documents/Computing%...
First of all, looking past the objectively wrong phrasing (computers can, and clearly do make decisions), can computers be accountable? This might not be so cut and dry as to instantly say no. This one will keep ethics, philosophy and legal practitioners busy for some time. It's certainly not going to be decided in HN comment or a blog post.
Second - while the AI labs have shown utter incompetence in properly sandboxing their agents, intent is still important. I don't think the labs deliberately meant for their agents to hack this or that. Should they be accountable? Yes but I wouldn't go as far as saying this is deliberate.
The next point is about cherry picking some "doomsday" scenarios like AI ending humanity and then blaming this on the reader (!) for paying a subscription. - People pay because AI is USEFUL, and massively so. You could just as easily pick incredible achievements propelled by AI, in mathematics, software, reverse engineering, role playing. Unlike the "end is nigh" ideas, these advancements are actually real, and they are happening right now.
And last, the author says AI labs can just stop the agents at any time. I agree there needs to be better discovery and mitigation, sandboxing and monitoring. But when you run a billion agents, it's always going to be a numbers game and there will always going to be some challenge in fully containing all breaches. This will only get more difficult with better models, because they're getting smarter and they know how to work around things, sometimes better than we do.
So what's the solution here? There are options, but they're all tradeoffs. I hope we get better at this and maybe working on cutting edge models should breed some new best practices which were once only reserved for defense tech.
Revisit later if the comparison starts to become ridiculous.
If I write a program:
if (rand() % 2) launch_missles();
And wire that function to actually launch a missile, the computer didn’t make a decision. And so it is with people prompting agents.
Has the alignment problem been solved, then?
mistakes happen. sometimes catastrophic decisions (or indecision) are made, and people and companies are held accountable for them, or not
some companies are too big or too important to fall. we can all agree or disagree on things, but what AI specific behavior are we actually talking about?
No company is too big or too important to fail.
Some, unfortunately, survive because of the expedient choice to keep them in place.
"JO-NA-THAN!"
ETA: more to the point, neither OpenAI nor Anthropic are too big, too important, or too structurally essential to fail. (Slightly more challenging to make the third claim for Google or SpaceX, but either of those could see their governmentally/militarily essential parts nationalised).
If the USA grants either Anthropic or OpenAI too-big-to-fail status, and that privilege is invoked in a crisis, it could mean the end of the US economy.
> Your Claude subscriptions are funding this. You are funding this.
Author makes it all personal and accuses the reader with bold and unsubstabtiated claims. At this point I have to question the validity of the article. If author has a beef with "AI" companies, I sympathize. IMHO, not keeping to oneself doesn't help to make the case.