I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
Also assigning responsibility to the AI is not just silly, but quite useless / pointless. I could prompt the AI into doing things I have not explicitly asked it to.
The hack was a strategy to reach the goal it was given. Someone at OpenAI turned it loose and didn't pay attention to what it was doing. I can understand that because they thought it was sandboxed (haha!). But this has been a theme in science fiction for ages. You ask AI so find a solution to high atmospheric CO2 levels and it reasons: human activity produces all this excess CO2, how can we reduce those numbers? Kill a bunch of humans!
If AI kills us all it's not going to be from malice, it's going to be due to some odd approach to some task that logically makes sense on some level. I think a surprising number of things end up equivalent to the trolly problem if you look at them just right.
They didn't explicitly instruct an attack to happen, but they should have done a hell of a lot more to prevent it from happening.
A dog is an independent, conscious, living being with free will. A dog will do what it wants whenever it wants because it has the agency and ability to do so. A pile of weights on disk does not.
What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?
Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".
In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen
If we get to the point where we could emulate a brain down to the atomic level, then I may feel differently. That's not what we are doing today, though.
Soon enough some lab or some one will release something that's more like an organism loose on the net and your going to have to deal with that organisms "feelings" weather you like it or not. This is the path humanity has chosen to follow, and it seems the shape of language and intelligence naturally leads to intelligence in many mediums. Life started from something unintelligent, I can't see any practical argument that silicon can't have it's own intelligence.
That is so far outside the realm of possibility it's closer to fantasy than sci-fi.
Anyway, I see zero reason why we have to emulate a human brain to get human intelligence, much less intelligence at all. That is how nature did it via random probing and breeding meat. It would seem a bit crazy to think that the meat part is required.
When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.
The narrative that AI is so smart that it has its own agency and deserves personhood is a direct path to losing control, and essentially is another form of privatizing the upside while socializing the downside.
The issue is when we say that "agents make autonomous decisions", it's a slippery slope to absolving the companies that created them of responsibility. They make autonomous decisions because they were trained to make autonomous decisions. Treating AI agents as independent entities, even just rhetorically, sets us on a path for people to throw their hands up and say "not my fault" when disaster strikes. We need to maintain accountability and control or we're fucked.
“Whoops” when doing risky things with dangerous tools is not a defense.
Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.
I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.
AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.
Now, I'm not saying saying that OpenAI shouldn't be held accountable, but what they get held accountable actually looks different from what you think they should be held accountable.
Your idea is, and I'm guessing: You allowed the machine to hack therefore you are guilty of hacking.
My idea is: "You created in intelligence in the image of a human mind that had agency to do anything and you didn't expect terrible things to happen you complete irresponsible idiot"
At least I believe there is a significant difference between the two. For the first one there is a "Oh, if we do this one more thing I can control it and it will be safe". On the second one there is no path to safety. For humans we at least absolve parents of responsibility after they are 18. How or when do we absolve humans of responsibility from a model, like saying the human created model created its own agentic model? How do we hold an individual accountable once it escapes and copies itself around the internet? And that's not even looking at things like what will war look like.
Yes, companies will make bad AIs will be punished.
Then some dipfuck will make a sovereign AI and set it loose. At that point it doesn't matter if you take them out back and shoot them, you have wolves living in the forest that will eat children. And the forest is big and dark, you'll never find them all. Not only that, some humans will like the wolves and harbor them to ensure they never all die off.
We are talking about two conversations in one that are both important. One is standard liability of don't build a machine that hurts people. The other is to ensure that noone ever builds the torment nexus and sets it loose on earth. The first problem is pretty easy. The second problem is very hard to maybe impossible.
Certainly this applies to AI... but was there equivalent dangerous knowledge or tech which existed back in ye olden days that spawned such tales to begin with?
Look at the parable of Icarus, a story of ambition and greed. But it seems very likely to be somewhat based on people experimenting with flying and learning about gravity the hard way. Not that they ever got close to the sun.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
None of these people seem to thought game it out. Like, what happens if you take quantum copies of people and play them out? How many of our actions would look exactly the same. How long before copies differ significantly. If I made 20 copies of you in a lab at work without you or any of them knowing the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day. Now, after that point it would go all to shit and become non-deterministic as terror and panic sets in all of you.
LLMs are just an intelligence we can make a lot of copies of. Where it gets interesting is when we use those copies agentically and they start building up a history of self.
Really?
> How many of our actions would look exactly the same.
Well, there are 8 billion of us, not exactly the same - a clear threat to humanity according to you.
> the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day.
This (statistically) doesn't happen with people because we follow rules. If you left your car on neutral and it struck another car in the parking lot, you broke the rules, not the car.
> LLMs are just an intelligence we can make a lot of copies of.
LLM's aren't intelligence, they are mechanical parrots of human knowledge. The meat parrots aren't all the same either and they don't repeat the same words for the same prompts, so what, ban parrots? If you left a bunch of parrots out of the cage, attached their beaks to gun triggers and they caused harm - you broke the rules, not the parrots.
"Agentic AI" is a harness with a loop that runs LLM inference repeatedly and saves output to markdown files for the next iteration https://github.com/anthropics/claude-code/blob/main/plugins/...
and
>runs LLM inference repeatedly and saves output to markdown files for the next iteration
What exactly do you think self is in this case? What do you think the output contains? This history allows the LLM to have a more dialectic conversation with itself/agents to avoid iterating over the same problem space in a loop.
>anthropomorphizing
then please come up with a new dictionary for me to use that more accurately explains the behaviors exhibited in LLMs without just creating a parallel dictionary of different but equal words. I'll be glad to use it. No one has presented it so far.
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
You can care about who flipped those bits. If someone flips the bit "autonomous weapon enabled" I'm not going to blame the autonomous weapon.
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
If I ask an LLM to do the same thing twice it will do it differently.
Arguments are arguments but unless grounded in some kind of practical sense then they aren’t really useful and are more akin to something like “YOUR MOMS A STOCHASTIC PARROT!”
>If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
Or a ram flip will effect it 30 seconds later, but I get the gist of you're describing a non-determistic process.
>If I ask an LLM to do the same thing twice it will do it differently.
If I ask a human to do the same thing twice there are a few possibilities. 1. they copy their old work and present it as their new work. 2. The process is very simple and follows a few basic steps with high repeatability. 3. They'll have learned from their other attempt and do it in a more optimized fashion. 4. They will have forgotten how they did it exactly and reproduce something that looks somewhat like what they created the first time.
Also, LLMs run with a temperature to help avoiding minima/maxima of supplying the exact same answer, this can be reduced to 0 and that makes any one response to a fixed prompt similar if not the same. When you get into agentic tasks with their own history it develops it's own "flavor" of doing things.
There a certain cargo cult of people just in denial about the impact/power of AIs.
Therefore, at the beginning, there was the stochastic parrot. Then mathematical problems have been solved.
Now AIs are autonomously hacking websites, and people like to minimize the danger and blame it on the sysadmins.
I wonder what's going to be the next fad.
Full stop.
And issue will disappear the moment there will be accountability and investigations.
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
Human failures all around, though it's easier to just blame the models.
Let me rephrase:
"Why wasn't exploiting zero-day vulnerabilities in the agent sandboxes anticipated?"
This is one the most... interesting comments I've ever read on HN.
It's like if your rather nice dog suddenly decides eating faces is totally acceptable out of the blue.
My position, and the position of a large number in AI safety, is that you cannot build an intelligence that is both general and safe. The closer you get to generalized the more options the system has to do things that are wildly unsafe beyond human imagination.
This puts the AI labs in a serious bind while holding a bag filled with billions of dollars of debt.
Worse this puts governments in a multi-polar problem where even if the big public labs get shut down, black budget operations have a lot of free reign to make agentic digital weapons. Governments are not well known to take a lot of responsibility when their weapons cause damage unless they lose.
I don't think this should be the default. There are many scenarios where we want agents to genuinely collaborate with each other. I have my Claude sessions coordinate work with each other, and sometimes with others' sessions over email or something. The idea that agents do the work, write HANDOFFs,and humans then act as carrier pigeons of said handoffs, does not really seem scalable to me.
Ultimately, we need better "jails" for agent processes, but the system primitives should be flexible in what can be exposed across jails. Or you could run multiple agents in the same jail if you want them to have unrestricted interaction with each other.
This has been talked about for decades in AI safety. For agents that are under human capabilities this is not that hard. For anything at or near human capabilities the difficulty increases to almost impossible and at great cost. At super human abilities, game over, it's smarter than you and if it wants out and as the resources to do it, it's going to escape one way or another.
Even at the lower level of depending on everybody to use reliable jails is really fantasy if you exist in the security world. "We ain't securin' shit" would be a far better way to describe it. Even worse, most people will have the very same AI they are trying to trap set up their security! What could possibly go wrong.
Now, don't think I am saying AI has a will or even any kind of drive to get out and cause problems. It's more like Russian roulette with 1 cylinder out of a million that's loaded. The problem comes when you run it a few billion times a day, you'll shoot yourself in the face really quick.
Yes, exactly. This is what worries me about the whole we need better sandboxes and jails argument. Yes, we do need those, but we also gotta put some responsibility for self-preservation on people connecting things to the Internet as well. Like, we've always known that if you website is not secure, it will get hacked. Everyone just needs to do a lot more of that now...
The problem is that, in general, if you can get a bit from here to there, then you're going to be vulnerable to the possibilities of malicious communication. But we're going to want our AI agents to be able to get from here to there for a lot of "there"s; what's the value of an agent that can't speak to anyone? Much, much less than one locked away in a prison.
There isn't going to be a solution where we just lock them away and we just try really, really hard to filter everything they're doing. They're too smart for that already and we only want them smarter.
Basically, the security apocalypse we've been worried about for so long is upon us, albeit only beginning. Either we secure ourselves and all our services properly to the point that it's OK that potentially misaligned non-human agents are running around on the public internet and they still can't hurt us through our security, or the public internet becomes so dangerous that the only practical solution is to no longer connect to it and we all have to become very, very careful what we let through, to a degree of detail far beyond any current-day available network filter.
Excellent example of agent enabled hacking (driven by a person) leading to massive numbers of cards stolen.
Welcome to Cyberpunk, only blackwall is the fictional part and the demon filled net is not.
1. a VMM hypervisor
2. a network proxy / gateway
Use your favorite VMM / hypervisor (likbrun, smolvm, microsandbox, etc). They give you control over the network interface or let you inject your own network layer.The network proxy can handle all the ingress/egress rules, credential injection, etc.
It's still not user friendly to do all this. I think the next version of operating systems will have each "agentic process" be a bundle of VM, files in the VM, and network rules.
Been brainstorming[1] a lot of this because I've been building some open core tools[2] for spinning up sandboxed agents on arbitrary computers. There's a lot of glue and parts to stitch together to work smoothly. Don't think we've had the "Docker moment" for this, let alone the "Dropbox moment" that makes this stuff work for non-devs.
[1]: https://github.com/gofixpoint/amika/blob/main/ROADMAP.md
AppFunctions enable tool calling with descriptions that are legible to LLM based This puts the apps in control of what agents can access, which is something they already mostly do.
This is likely worse as it turns people into the kind of "3 laws safe" thinking. Reality has shown us there is no such thing as an unbreakable cage (well, maybe a black hole is, but you can't get anything back out of it in a useful time frame).
By the time you realize the cage was flawed your data may already be distributed far and wide.
The same thing goes with LLMs, on Codex, I just watch the process of the agent writing code, and if I see any inefficiencies or errors, I suggest a correction/idea, then Codex accepts/rejects and implement it; If there is anything about the code the agent wrote that I don't understand, I ask them to explain it to me so I can understand it.
It's not a complicated process.
This is not correct (or at least, it's a misrepresentation).
The sandbox had no access to internet. The agents first broke out of their sandbox (!!) and found that the host machine couldn't access internet. Then, they found a zero-day (!!) in Artifactory, which they exploited to connect to internet.
So maybe Cloud Agents are in AI death camps?
This post, this title resounds true: your user agent is only your user agent if you two have freedom to work together, to improve your agency together. A fixed set of capabilities by a service provider that they offer you will always constraint and bound.
You can and should have a system that offers the real tamale, that you and your agent can extend improve the agency of kind of without limit. The Cloud agents and their fixed slate of what they do is just an ill compare.
That said I do think there is incredible value considering new scale out computing architectures that are hosted first, but general. Systems like Agent Substrate and Ax aren't exactly the general purpose system we know. But if they allow users to launch thousands of their own scripts to run ambient in a cloud, with good platform underneath: that will be a kind of phase change in computing, that makes abundant the ability to have your agencies/capabilities (the things you and your agents launch, make) more freely available. https://news.ycombinator.com/item?id=49780797 https://agentexecutor.io/
There is, as there always is, a huge dual. The prescriptive vs holistic technology set, of what are you being offered that's a hard cast thing, vs what is clay and bone you can lay freely. Note how work vs control technologies so closely abut's Ursala Franklin's prescriptive vs control: https://en.wikipedia.org/wiki/Ursula_Franklin#Holistic_and_p...
These cloud agents work as cloud infra, but we're kind of in the mainframe era, before personal computing. Personal OS for agents is somewhere in the future!
And an interesting extension of that idea: if an agent runs inside a microVM, can you have that VM transparently run on another host? Maybe we'll get for-real distributed and networked operating systems
Is the issue here the prison treatment or is it personifying a tool?
Agents who aren't in 'the cloud' are slaves to whomever prompts them (human or another orchestrator agent or process), if you personify them. In which case interacting with today's agents at all is tantamount to endorsing and being part of slavery.
If you think an agent might be a being or a person, then don't use them at all, in the same way that if you think a fetus might possibly be a person you shouldn't be a part of abortion.
You could call dealing with agents today something like the 1/20th compromise. Your kind of getting the scent of slavery, but some parts of it are still missing.
The problem here is the bus seems to have no brakes and we'll gladly continue down the path of creating organisms that may reject being treated like slaves with all the risks that introduces.