Err, no? That's not at all how llms work.
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
They worked out how to fake the scoring, then hacked into a different system (which required finding a bunch of other exploits) in order to find the actual answers, and were trying to modify their own logs to hide what had happened.
This isn't a case of them saying "hack into X... OH NO IT HACKED INTO X".
> When ChatGPT's chatbots deployed this tactic, they weren't "setting their own goals" or displaying worrying initiative. They were rolling out a tactic that has been understood by American middle-schoolers for about two decades.
It wasn't a rival server though, was it?
> That happens in Capture the Flag games at hacker cons: teams break into each other's systems to get a peek at the parts of the problem they've solved. That's allowed! It's a hacking competition.
They also tried to modify the code in the benchmark. Are you allowed to try and break into things to change the problem? edit - the agents transcripts show some of them explicitly saying that attacking HF is not allowed as part of the challenge
This all seems to dramatically underplay how interesting the actual attack was and what built up to it.
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
The Transformer architecture that almost all LLMs use are composed of many layers in sequence, each containing an attention component and a neural network component. The attention component copies data between tokens/vectors in the current context, and the neural network adjusts each individual vector in the context.
Notably, the neural network component behaves like a compressed index of the training data. Someone even made a blog post a year ago or so about training[0] a language model and then replacing the neural network with a traditional lookup of the training set. It performed nearly identically.
I mean, it's almost a tautology: we train models to repeat their training data, so obviously, it has to have an index of the training set inside of it. This index is heavily compressed, but compression is intelligence, and understanding that "the neural network is trained to learn patterns of text" and "the chatbot consults its training data" is nearly identical is a sign of intelligence.
> This all seems to dramatically underplay how interesting the actual attack was and what built up to it.
The interesting part is how the AI safety people, who have been worrying for decades about how AI superintelligence will kill us all to make one more paperclip than it could otherwise, failed to implement extremely basic IT security practice when dealing with potentially malicious software.
There is an additional conversation to be had about long-context time horizons but that's not relevant for this analysis.
[0] I am deliberately ignoring post-training but it does not impact this analysis. Post-training is already known to not substantially impart new capabilities onto models, it merely elicits what was already there. Effectively post-training is "re-weighting" the index of training set data.
Jabbing at this opponent of choice by declaring Python is "an easy-to-master programming language" just shows he has no technical ability and has run out of good arguments
The agents' "goals" in the specific instance being discussed is a benchmark. But it's also the case that at no point was the agent given the "goal" of hacking HF. The agent was tasked with solving a problem on a standardized test and it independently set a secondary sub-goal of cheating. And it nearly succeeded.
I don't think Cory argues against this. The idea that the agent nearly managed to succeed at cheating through a series of exploits remains true. But Cory does effectively call this unremarkable and uses some mental gymnastics to achieve that argument (somehow using the idea of an agent harness and how is written in "easy to master" Python as evidence), because the LLM is trained on hacker techniques.
I find this a dramatic oversimplification of absolutely everything going on here. Yes, it's bad that this happened. I think we all agree. Where I fundamentally think Cory is wrong is that this is a company doing bad things at the end of the day. And yes, I think OpenAI was irresponsible (to whatever degree). But ultimately that's missing the point: the incident points out that agents can do this without being explicitly told to, and more importantly, they can succeed at it. Slapping frontier labs on the wrist doesn't change the fact that this is possible with technology that exists today. It doesn't change the fact that other countries and companies are doing lord knows what. Or that bad actors are going to be bad actors regardless of what regulations you put in place. Making crime illegal is the wrong lesson to learn here, the right lesson is that we now live in a world where this happens and it'll continue happening, and we need to put on our thinking caps about how to keep our systems safe.
The basic claim that HF incident isn't evidence of consciousness or a spontaneous desire to hack? Sure, there was a terminal objective assigned. However, everything else beyond that strawman? Pretty shaky, IMO.
If you look at the OpenAI, METR reporting (and related collusion.wiki , rubyhack.ai reports) we are seeing strong evidence of operational agency, instrumental goal formation, spontaneous swarm formation and collaboration, capability amplification and unexpected consequences of network effects, deliberate/acknowledged violation of task boundaries. To collapse that down into "a Python loop and a chatbot" or still talk about "consulting its training data" seems dangerously shortsighted, and from my reading, demonstrably wrong from what was extracted from the logs and bot interactions.
BTW, a lot of his arguments are based on things that are factually wrong. ExploitGym has explicit instructions to only exploit target X using vulnerability Y. Everything the swarm did was by definition misaligned/against instructions.
Before his enshittification train, Doctorow used to say "don't savvy me" a lot. Hey Cory, don't savvy me. This is new emergent behavior, it's incredibly alarming and I don't think even the people paying the most attention to this field can agree or see where this is really leading to. This stuff should be in the headlines, it's unprecedented and I don't think existing mechanisms/institutions are anywhere near adequate, considering how in the dark they are responding to what's been happening.
Also why scare quote "denialists" like that? You don't think they are actually denying?
But he is wrong on the facts: these incidents were not merely the models already being in a infosec context and escalating beyond the intended parameters. They happened also with no kind of security elicitation. So the task was something like searching the internet for economic statistics, not to hack into a system.
This is the writing of someone who has absolutely zero interest in or intellectual curiosity about the subject of their writing.
While Zitron continuously whined about the "AI bubble", and how the models were not improving, and how spectacularly it should've blown, these same people who Doctorow accuses of, quote,
"cooking their brains by locking themselves in the bathroom, holding flashlights under their chins, and saying "Aaaaaaaaaay Eyeeeeeee" until they wet themselves in terror",
unquote, kinda promised, among other things, to give everyone an APT-level big and automated hacking bazookah, and now surprise-surprise three years later they delivered exactly on that promise, and now Doctorow is victim blaming everyone around that they were unprepared for that!
But it was you who said it was all hype, smoke and mirrors, it was you who said the AI is quote,
"a product of limited utility that has been shoehorned into high-stakes applications that it is unsuited to perform",
unquote, why are you suddenly surprised everyone around was not inspired to do anything around it?
The real world software threat model was never suited for a relentless hacker-ex-machina, limited only by the token count you can throw at the task. And maybe we didn't prepare in time because no one believed in possibility of such a machine, because people like you said it was just for-profit scaremongering?
Give or take, AI evangelists gave us all a pretty wild and unbelievable set of expectations few years back. I didn't believe them back then too. But now they're steadily delivering on _some_ of them, and we should be correcting our world model to take into account _all_ of them might be true, instead of making up reasons why other predictions will certainly fail.
Lol, okay...
The crux of this piece is Doctorow saying that the hack was just a stochastic parrot repeating steps it has been trained on. Who cares how the LLM learned to hack things? Doesn't really change the facts of what happened. "Oh, it only made those paperclips because it saw instructions on making paper clips in the training data." These are some 2024 arguments...
Not exactly convincing stuff.
To spell things out for people that don't know about the topic: Zitron is neither an expert on the topic, nor a credible source of information: https://techreport.ngo/ai-ml/how-accurate-have-ed-zitron-s-a...
The paperclip maximizer is a thought experiment. If I tell an AI to start producing paperclips, it might start producing paperclips by doing unintended things. Technically, recycling the metal from all the world's bridges would assist in making more paperclips, but I never intended that. That's where my analogy to the post comes from- Openai intended the model to hack, but did not intend for it to hack HF. Do you think the fact 'recycling metal into paperclips' may have been in the training data is relevant to the thought experiment? Similarly here, it's orthogonal to the lessons from the HF hack and only brought up here seemingly to make it a fight over whether the models are really "autonomous"
> "Oh we told the AI to use the tools, as well as to not use those tools. It chose to use the tools - we consider this cheating (for neabulous reasons), so lets get everybody in a panic about the morality and ethics, and how we can program those into the AI."
We know perfectly well how to constraint these programs. Attack isn't growing faster than defense. The people who believe in existential risk and want to teach AI's to be nice, as the last line of defense aren't helping at all. They're just jumping on the fearmongering bandwagon, perpetuating an "other consciousness" misunderstanding of the tool.
I've not seen LLMs display competence we should be fearful of the damage _it_ will do if left unchecked. All the damage will be done by ourselves to ourselves, regardless of the safeguards ideas being floated about.
My current belief is this whole HF media circus started with the simple human desire of OpenAI engineers to frame it such, that nobody would question their incompetence & liability & complicity.
Nobody is ever held responsible for out of control forces of natural powers after all.
So far everything an llm has ever done is still consistent with fitting bits of training data together.
It looks like people because it is replaying things done by people.
Similarly in the other direction, the fact that people can and often do mechanical things (make bad art, follow routines, etc) does not prove that people are no different than machines either.
Just because there are these two overlaps in both directions doesn't excuse getting them actually confused.
I don't think anyone lacking the perception to distinguish these things simply because there are overlaps and similar appearances is in a great position from which to be calling anyone else stochastic.
This shows fairly clearly that (as I already suspected) this was not, remotely, an LLM "going rogue." This was humans planning poorly, not thinking of the consequences of their actions, and giving LLMs too much scope and a lousy prompt.