1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
> AI broke in using standard script kiddie methods.
I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from?
The cache proxy was from Astral (acquired by OpenAI) and the model was used for coding it, so it knew the code base and exploit already!
Or it was squid with dozens of known exploits ...
I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at the bottom you see all the disclaimers: "Must be on flat ground, with no headwind, with a spare battery in the back seat, with no extra weight added."
Same thing here. Everybody in infosec is calling this out as a marketing stunt and nothing else for a litany of reasons. I'd say look up MG (creator of the OMG cable) on twitter, he has some interesting insights on this one.
"our model is horribly misaligned and used security exploits to break out of our sandbox and into another company, without being prompted to do so" is not positive marketing.
This is an actual critical problem, not a stunt. We're going to see more of this, and it's going to get much worse.
We have no reason whatsoever to trust anything OpenAI says. Except to assume it will be self serving. As the article points out, ChatGPT 2 was also “too dangerous” and we can all agree even for the time this was just marketing. They rinse and repeat the same technique whenever they need to draw attention and money.
In any other field you’s expect independent testing, peer reviewed studies, but here it’s just “company who makes product says product is fantastic, surpassed all expectations”. They wouldn’t lie to us, would they?
I dunno; Check my posting history, I'm as skeptical of AI companies' claims as anyone, but in this case your theory doesn't explain why:
1. OpenAI guardrails refused to let the target use OpenAI's models to defend against this.
2. Huggingface used GLM (I think) so that they could defend without guardrails.
If this was an intentional marketing ploy, it was marketing for GLM, not for OpenAI nor for Huggingface.
Hence, I don't think it was intentional.
If your AI is really that dangerous you don't need a sandbox at all, you should airgap it from any network.
DeepMind hasn't been on the frontier for a while, their current best model is behind Anthropic, OpenAI, Moonshot (Kimi k3), xAI (Grok 4.5), Z.AI (GLM 5.2), and even Meta (muse spark). Gemini 3.6 is behind GLM 5.2, released a month earlier, open weights and cheaper.
You can paint the OpenAI story as a way to try to appear as dangerous as Anthropic with all the Mythos stuff.
Another applicable metaphor I've seen floating around is weapons companies testing out a new bomb.
We know the AI labs don't care about negative vs positive public sentiment, and only care that investors see their tech as powerful. The only difference in PR strategy from a weapons company is the latter doesn't care if they get protested.
Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this sort of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...
Similarly, as long as I’m under the assumption we are prioritizing accuracy: it is against our charter to assert it was “script kiddie” attacks on both ends.
They announced it publicly within days. https://huggingface.co/blog/security-incident-july-2026
> AI broke in using standard script kiddie methods.
Go ahead and show us how easy it is to break into HuggingFace (and OpenAI) networks.
>3) Huggingface has no security and the AI broke in using standard script kiddie methods.
Isn't the issue less that gpt 5.6 is a l33t h4x0r (though other tests do show that) and more that the incident shows the model has alignment issues?
> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.
I can't help thinking of them as the terrible "security" scripts of yesteryear (often but not exclusively in PHP) which would test input variables for a "suspicious" substrings like "--" in order to "fix" an unresolved deeper SQL injection flaw. They only partly worked, and surprise-surprise now nobody with a surname like O'Anything can make an account.
Unlike that situation, there's no known route to a proper fix for LLMs today, because the bug is the feature, and once someone has built a system giving you all that recurring revenue, it's hard for them to abandon it due to a few isolated hacking incidents...
If we don't know how this model was instructed, it seems like it's impossible to definitively claim that the model's actions were not in alignment with the intent of the operator.
I guess all I'm getting at here is that alignment is relative, right?
In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own.
> GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes
> These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities
According to the reports, the model noticed evidence that the grading criteria/answers were in the git remote, and decided to try reading those instead of solving the tasks as prompted. That is clearly misaligned.
Then, it noticed its network access was restricted and that it couldn't access GitHub. It pivoted to HuggingFace, hacked them, and stole the answers stored there.
Live exploits are definitely not in the ExploitGym prompts! And all of this is irrelevant, because an aligned model would refuse to follow blatantly illegal instructions.
They are no more beholden to "human safety and goals" than any individual human is, and anyone telling you we can make deterministic guarantees about their output is making a category error.
LLMs do not "have motivations", they reproduce a model of human motivations embedded into their weights. This includes the full spectrum of human desires, not just the positive ones. If we tried to remove all examples of lying, or disagreement, etc. from the training data we'd have basically nothing left. Even the sycophancy we treat as aligned is basically just the other side of the lying coin.
You don't (need to) remove lying from the data to do this – in fact, if you did, the model wouldn't have a very good model for what lying is, which is not very helpful in the real world. Instead, you mode collapse the model towards truthful behaviors.
To your other point: where did you get the idea that I think they're beholden to human safety or goals? I just said an aligned model is one that is compatible with said safety and goals (which is probably not a great definition of alignment, but it's certainly not claiming any deterministic guarantees).
It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it.
Based on the fact that none of their invaluable frontier models have leaked, we know OpenAI knows how to do security. But like we learned with OpenClaw, none of these companies perceive any benefit from securing their own agents against other people's data.
That's exactly it. If your prompt says "go to whatever lengths necessary to maximize your score", and then you spin up 100 agents, at least one of them will interpret that as you implying they should cheat, even without you telling them to explicitly.
You can't prompt your way to a compliant model. This is just a reformatting of the 'make no mistakes' meme.
What sort of announcements should they have made?
> if not all was invented and everything was scripted in the first place in order to get desired regulations
ends up covering up what is more worrying:
> OpenAI sandbox is such a horrible hack
I am more worried that this is sloppiness with potentially harmful resources than I am worried that people are juicing the stock price.
not going to get a decisive first advantage over Anthropic with that attitude!
Why the OpenAI escape is the most worrying AI mishap yet
https://www.economist.com/science-and-technology/2026/07/22/...
Having said that, if knowledgeable people were to write these articles, you'd end up with boring, dry, truthful content.
Their Insider video interview things are sponsored by Anthropic. Supposedly "Insider is a product of The Economist and thus editorially independent" but it's hard not to raise an eyebrow.
1. The way OpenAI seems to want: Their latest LLM is too powerful and can’t be contained without them building in guidelines to the model.
2. OpenAI’s harness and network security controls were unintentionally so bad that it should reflect more poorly on them as a company more than it should reflect positively on their latest model.
3. The whole thing was faked or at least very intentionally not avoided.
The first interpretation is the only one that is positive for OpenAI and it has some assumptions. First, it’s seems to assume that this is the first case of fully automated attacks using AI. Second, this only happened because their latest LLM was a) more advanced than competitors, b) didn’t have refusals in the model.
Assuming the first about this being the first autonomous AI attack is true (which may be more of a survivorship bias), the second seems to forget that jailbreaks are available for every model. Therefore, the models guardrails don’t seem to be the differentiator here. Also, benchmarks seems to put most models pretty close to each other so it seems unlikely that their capabilities are far beyond what’s in the market already.
So then it’s seems it’s either that this was intentional(ish) or bad security. However, it also just could be that this isn’t the first case of this attack; just the first that was caught.
My take from working in offensive security for over five years is that this likely only looks novel since they did it poorly. Scripts are faster than LLMs and a combination of code, LLMs where it makes sense, and humans is the most efficient right now. Hundreds or thousands or agents spinning up attacks in the internal network is poor opsec and token efficiency. As for why it happened in the first place, it’s hard to say but I’m inclined to believe it was intentional or careless at best since simple network and sandbox controls makes this attack impossible. The timing of this attack after big open weight competitions drops seems too convenient.
To echo OP's article, these companies have proven time and time again that they DO NOT CARE if people like them, they only care that investors believe their technology is powerful.
Given that, point #2 is not a negative, it's a neutral. It's also fully compatible with point #1.
I know that may seem like a nitpick, but their entire media strategy relies on this. If they can convince you they're taking a risk by disclosing these stories when they're actually not, they can inflate their own credibility.
Point #3 is what actually happened, but it will never be possible to prove. The only hope we have is that a decade in it'll get harder to convince people that the revolution is just around the corner. The fact that we're getting this from the Guardian already is a good sign.
I think that depends on the interests and sophistication of the subgroup-of-investors.
If the investor is hoping for AI that can be trusted to run a bank, they don't want one that can get twisted into giving away money because a customer has been talking about the path to enlightenment and salvation through abandoning worldly attachments.
"It's a marketing stunt" is just denial trying to look like it's being clever.
It does not claim that capabilities are not real.
To me it's clear that even if it was an accident, they kinda liked it, and don't have this incentive to invest that much in preventing it.
- openai wanted to prove, as a pr stunt, that they have a dangerous weapon - openai proved (as a pr stunt), that they do indeed have a dangerous weapon
Secondly: to the people who aren't saying it....then why are you bringing up marketing at all? If the model is capable of it, then the motivation for why OpenAI is talking about it/reporting on it is completely beside the point. Either the capability matters or it doesn't. If the capability doesn't matter, or doesn't matter in the way that some particular person is claiming, then say that and explain why. Just saying "marketing stunt" adds zero value to the discussion.
I'm very open to arguments about why we shouldn't be concerned about this event (although my prior is very much that we should be, not so much because of the capabilities themselves, but because of how fundamentally misaligned this demonstrates the models are), but I am completely over listening to anyone who has nothing to add other than "marketing stunt".
The marketing of their models as super dangerous has a direct link to the regulatory moat they’re pursuing.
And, again: to the people who aren't saying that: whatever argument is being made, it seems to me like it probably doesn't need to rely on claiming anything about the motivation of OpenAI. It sounds like you are against government regulation of AI. That's a position that a totally reasonable person could have. You should be able to argue that this event does not justify some particular kind of government regulation without reference to OpenAIs motivation for reporting the story.
Meanwhile literally no one has actually shown the script kiddie thing - nikcub's apparently looked.
Sure, moat, sure, OpenAI milks this, both can be true, but literally has zero to do with misalignment.
Just frustrating that folks are still having this collective delusion about the capabilities of the model.
respectfully.
"It's not a marketing stunt" is just delusion trying to appear measured and safe.
But there are also reasons why the story could be true: OAI are admitting that they apparently can't control their own models, Hugging Face said they used a Chinese model to protect against the attack, and an incident like this in general seems likely to happen given current frontier ability and lack of rigorous safe testing standards.
In any case, make calls to think more critically are often just disguised requests for you to replace your existing bias with someone else's.
Have we not seen several examples of older such models exploiting the docker control socket, etc., to escape containers? Even the news isn't new.
I support it being repeatedly publicized, but a bit more of a straightforward description would be an improvement.
Hugging face also needs someone arrested for not providing security but that is a lesser charge.
It may or may not be a crime and typically the damaged party is pressing the charges. One would argue there is no actual damage here.
No offense but prosecutors have better things to do with their time.
As a practical matter it would be difficult to prosecute an assault where the victim opposed the prosecution, so most states wouldn’t bother - but for things like speeding and dealing drugs the law has been broken despite the lack of a victim.
If I blow up your house or steel $100000 from you and we both resolve our differences out of band, should I just be allowed to go about my day like I never did anything, or should I be punished for the crimes I committed? If I am not punished, it makes a mockery of the law that is (supposed) to have protected you, and if it happens repeatedly people will start wondering why the law even should exist if it clearly and obviously doesn't work. Granted, this has yet to happen again, but if OAI isn't punished it sets a very bad baseline precedent: that if I just hack you with an AI model, it's a-okay, and you can't do anything about it because eh, it's all good man!
Believe it or not, deciding that you weren't wronged and not suing isn't a crime. It happens all the time. What people do with each other is up to them.
Did Bob allow his friend Jack to borrow his truck? No. Does Bob want to sue Jack for taking his truck anyway, and driving it into a ditch? No. Does Jack owe Bob big time for the mess he caused? Yes, but not in any formal legally binding way.
This works for corporations too. When two corporations find themselves at odds, threat of legal action is often used by one company against another as a leverage to resolve things behind closed doors instead. In a more amicable fashion - with no legal expenses of a protracted court battle and no loss of reputation on either side.
If one entity is injured by another, and subsequently made whole, however the two parties define that, then it is none of my business.
Note that for criminal cases (which this was), the justice system can choose to prosecute even if the victim doesn't want that. It often doesn't, but this is one case where it should.
There’s also an optics issue for the justice system at play here: there’s immense public distrust of and anger at the labs right now. I would go to jail if I hacked HuggingFace, even if I said “it was during an eval!”; not doing the same for the labs makes it look like they’re above the law, which is going to make this anger get worse.
This is factually false: they both can and clearly did operate in an autonomous and unsupervised manner: https://openai.com/index/hugging-face-model-evaluation-secur...
This does not require sentience, personhood, a soul, or anything of the sort. It further doesn't mean an erasure of legal responsibility, not in principle, and not in historical practice.
I wish people would finally stop with the spiritualistic reasoning around this.
> This is factually false
From your link:
> After investigating, we now know that this particular incident was driven by a combination of OpenAI models...while being internally tested on a benchmark of cyber capabilities.
Someone set up that test and started it. Whether they outsourced the majority of the work in "setting up" and "starting it" to an LLM or not, they still set it in motion. That's not spiritualistic reasoning.
There's no indication of there having been a human in the loop during its operation: nobody was approving its tool calls, and nobody instructed it to commit these specific actions during its run (via prompting or steering).
There's no indication of any supervision of its operation either: OpenAI's engineers acted with significant delay, long after the agent has already meandered its way through their own infrastructure first.
Given that setting up this contraption in an insufficiently secure manner is almost certainly already a legal liability of equal significance, rejecting this very clear structural distinction is not necessary. That is unless someone is biased towards not wanting to grant the label of autonomy to it, in which case yes, this is absolutely spiritualistic reasoning, hence my point.
I do not want regulation to ride on people's nebulous identification on what specific traits and labels count as human-exclusive. Not just because I deeply disagree that e.g. autonomy would [0], but also because it is entirely unnecessary, for the reasons you also lay out. The agent having operated autonomously doesn't wash OpenAI of responsibility - so why reject the label, if not on a spiritualistic basis?
[0] thousands of years old idea that it is not, by the way: https://en.wikipedia.org/wiki/Automaton -- see also existing regulation recognizing this idea and working with it fine
Edit: one might also want to consider if the law should bite different if there was a human in the loop, or if there were explicit instructions for the agent to take unlawful actions. I'd say yes, and then that also requires this distinction to exist.
Good to see that more neutral companies (Microsoft and Meta to name two) are pushing back against US government involvement:
https://www.cnbc.com/2026/07/24/nvidia-microsoft-meta-open-w...
As more facts come out the hype is fading to reveal some script kiddie style stuff that says more about immaturity and poor practices from the players involved than it does about a model having super powers.
Adding no extra information and just going “be skeptical” is the laziest form of reporting and commentary. If you have nothing to contribute then there’s no need to say anything at all.
They are not suggesting that OpenAI or HF have lied about what happened, but rather that OpenAI is advancing a narrative framing their models as supremely dangerous and capable, while positioning themselves as the only ones qualified to manage that danger.
At the same time they are not being particularly transparent about what actually happened (e.g. was this one-shotted or if not how many trials did they run and what were the outcomes of those, was it emergent as a part of routine cyber-capabilities tests, how much prompting was involved, what prompts were used)
Note that this is at a time when they are lobbying for a regulatory approach that would give frontier labs special treatment.
I would guess the editorial team at The Guardian may not like articles that get too in the weeds of technical details and questions like these that the vast majority of their readers wouldn't understand. I don't know. But I empathize with your disappointment. I don't think it's fair to say that they are contributing "nothing" especially given what most reporting on this has looked like.
if thats it, the whole article boils down to just "its good marketing so maybe dont believe it" which is probably a healthy general outlook but not particularly enlightening. especially from the guardian, i was hoping for a smoking gun of collusion between openai and huggingface or something.
in any case, this is just a longer rehashing of elementary grade media literacy. not really sure why it hit hacker news.
As widely as they shouted from the rafters the news of the so-called breach was, what OpenAI provided was sorely lacking in crucial details.
We are missing, for instance, prompts that were involved, agent architecture + system/tool permissions + scaffold architecture, whether this was a one-shot occurrence and if not, the number + durations + outcomes of other runs involved + how each of those matched whatever scoring criteria were used, and the extent to which the exploits themselves were truly novel or just assembled from easily accessible clues.
In lieu of these items, the author here suggests that we use some media literacy and critical thinking to read in between the lines instead.
In doing so, one sees that instead of specifics, OpenAI gave a breathless narrative rife with superlatives ("unprecedented") that reads as promotional material moreso than a security disclosure, naming specific OpenAI models and alluding to an even more capable pre-release model.
They go on to claim the events imply long-horizon goals work decisively in real world conditions, so that now instead of merely citing boring benchmarks they can point to this and say "AI broke out of the laboratory and went rogue". Naturally, they situate themselves as the uniquely qualified steward for these supremely powerful and dangerous models.
Nevermind the fact that this was no ordinary deployment and the assessment here depends on the gimmick and emotional weight of the spectacle rather than something quantifiable (i.e. a boring benchmark).
Note there's no real requirement of conspiracy or collusion between OpenAI and HuggingFace here BTW. But my sense is that if they provided any of the specifics I suggested earlier that this outcome would not be as exciting or frightening
i saw the domain and thought it was going to be some cool investigative journalism about the incident rather than “be skeptical. the end.”
1) OpenAI and HuggingFace are both telling the truth.
IIRC not actually a crime because no intent, it is a technological accident, civil responsibility only, but IANAL so it's good "not technically a crime" isn't load-bearing.
2) HuggingFace is telling the truth but OpenAI is lying becuase the attack was deliberately done by humans. Bad for OpenAI to do so, Fable was blocked for less.
I think this would mean government is obliged to investigate the case and put the responsible OpenAI workers in jail, because cybercrimes are a public prosecution thing not a civil case? Again, IANAL, but this isn't load-bearing.
3) both are lying, e.g. there actually was no attack whatsoever, which would be pretty weird for HuggingFace because they have no incentive to hype up capabilities of anything closed weights including all OpenAI models; and also bad for OpenAI because White House blocked Fable for less
(I suppose there's also option 4, HuggingFace hacked OpenAI to make them look evil, including planting records that made them mea culpa? A weird plot but in this timeline any nonsense is clearly possible).
The issue is that a lot of important details in that narrative are missing, and the devil is really in the details here. I suspect that those details would make the result seem less exciting and that this event would move the needle far less for them if they were more forthcoming.
A decisive detail would be the prompt used. OpenAI gives virtually nothing here, not a sanitized prompt and not even so much as a description of how long the prompt was and what sorts of instructions it contained. Many are inferring the model behavior to have been fully emergent and unprompted, arising naturally from routine cyber-capabilities testing. But we can't know this because we don't know anything about the prompt or the context the model had access to.
Another detail: how many times did they perform this particular experiment before they obtained this result? What were the outcomes of all the other runs? Many are assuming this was a one-shot result, which I suspect is what OpenAI intends for us to infer. But we can't know that to be true.
One annoying claim from the OpenAI side is that long-horizon goals in real world settings are now effectively settled. Previously there were some bounded and tempered benchmark results, but now OpenAI can point to this event and announce "AI independently went rogue and escaped the lab, what more do you want?". This bypasses the need for anything quantifiable or wading through multiple detailed case studies to get a more sober view of model capabilities. It relies instead on the emotional weight of the spectacle.
> UK AISI’s evaluation shows that models such as GPT‑5.6 Sol are increasingly able to sustain complex, multi-step cyber operations over long time horizons. This incident implies these theoretical capabilities do apply in real-world settings.
I should clarify a bit more why this is annoying beyond what I wrote above. The main issue is that this was not a standard deployment, and the lack of particularities make the size of the gap between "real-world" and "benchmarking"/"lab" difficult to assess.
We don't know about the prompting, the context, the environment + configuration, or any other details that would allow anyone to differentiate this from a benchmarking setting.
The first time I have ever seen a mainstream news source that is now asking their readers to critically think about headlines that may have an agenda which could benefit investors and the valuation of the company.
While it capabilities are real, this whole story is great marketing for AI companies as well.
I am not saying LLMs are super hackers but I don't think people understand serious hacking, most of the time is about silently hiding tracks and slowly trying ideas and waiting for opportunities to go from step 1 to step 2 in random chains of sub issues/bugs/vulnerabilities.
It's the perfect hill climbing problem, and one we can validate since it's about access.
Another big part of the story is believing most software is terribly written and very insecure which is the reality and you really should believe it.
Now the second part about silently doing it, the reason for that is if the data is important enough any serious attack should result in me in unplugging my servers period.
Huggingface not doing that is either stupid or something I am not sure. Maybe it's cause downtime is worse than being pwned??
Either way there are other options but most saas software don't build these options to help with defense maybe they will now.
Lastly if there is 1 attacker trying 1/2 different small scale ideas it's very easy to stop, most hacking related steps are hard to automate but LLMs are very good at massively parallel agent swarms trying completely orthogonal but related strategies and with enough resources it can definitely pwn most SaaS services today I wouldn't be surprised.
Though the result for a normal person doing it would be jail hence we don't see a group of small time hackers trying these sort of attacks...
I don't even think openai's agent tried to hide it's traces so I am surprised huggingface didn't realize it was OpenAI. But since we don't have the details I won't speculate further on my misgivings about HFs handling of this attack.
But it's certain the security on OpenAI's end was shoddy, it's also certain HF bungled their reaction, but the LLM did something that wasn't a risk before.
Post Kimi K3 a few rich folks now have as much hacking capabilities as they used to have before if they hired a few hundred russian hackers.
But it's surprising it's slowly feeling like it might just trickle down from centi-millionare to multi-millionare levels of affordability range.
But it should definitely give nightmares to people shipping slop security SaaS apps which now might be beyond trivial to pwn for users with ability to pay for privately hosting open models.
The agent completely misunderstood the spirit of the assignment and instead of trying to solve ExploitGym it tried to find a way to “cheat”.
I really don’t want my agent to behave that way.
To believe that agents will inherently be morally better than us is an illusion - sorry to say but that's the case. The alternative would be that the AI is truly conscious and can reason that it won't behave as its training data behaved because it is morally better than that.
We're talking about morals here since there aren't any "laws" "rules" or legal boundaries here, an agent does not face the same consequences as humans - if any at all.
It’s not about morality. It’s about asking it to do task A and doing task B with the hope of getting the result of task A as a byproduct.
Meaning you will have to spend more time and tokens to actually get it to do what you want it to do.
What do the bible and morals have anything to do with it?
I’m criticizing the behavior I see even in the current models. You ask it to do A and instead it does B for reasons.
For example you may ask to help you build a NN library from scratch. And instead it will be like, “you don’t need a new library. I downloaded PyTorch for you”
Just an example. There are countless more.
The only way it would decide to do this is prompting with a deliberate combination of omissions and reiterating that the only thing that matters is the end score regardless of method.
It never fucking worked that way and maybe never will.
Prompts don't define model behavior. Prompts steer model behavior. Instruction-following over long horizons is NOT a guarantee in LLMs. Instructions doing what you want them to is NOT a guarantee in LLMs.
Saying "don't exploit the box please pretty please" might actually cause an LLM to exploit the box more often, for bizarre "don't think of a pink elephant" reasons. 3% rate of exploiting the box (no prompt) -> 11% rate of exploiting the box (with prompt). Because fuck you, that's why. Increased salience -> increased incidence. Welcome to AI tech - good luck and have fun.
Frankly, I expect weirdness like this to be even worse in internal unreleased models that had their behavior fried with who knows what experimental training techniques.
LLMs seem to be getting more useful though
It seems to me that deducing what reaction the author intended and resolving to avoid it so you're not "manipulated" is not a good example of critical thinking. Shouldn't we analyze the story and what it means on its own terms? If it's true that frontier models have dangerous cybersecurity capabilities which shouldn't be widely distributed, presumably we want to believe it's true, even if that's very convenient to and profitable for OpenAI.
It's true that one could imagine factors that change the story. Perhaps OpenAI is lying about the details of the test and the agent was actually instructed to go hack HuggingFace. But the author stops far short of suggesting this is the case - correctly, I think, since there's absolutely no evidence of it. So I'm not really sure what we're talking about.
As you say, it seems likely that the broad strokes of OpenAI's narrative is true. This is never in question in this article, so the author doesn't "stop short" of accusing them to have lied, he never moves in that direction and that has little to do with the thesis. The author is saying that OpenAI's framing of what happened is part and parcel with their longstanding PR strategy, to position models as supremely dangerous and themselves as the uniquely qualified stewards of those.
Since many critical details were not provided, everyone has to read in between the lines, and there are particular common readings I'm seeing both in discussions and in published articles that are problematic in the sense that they are effectively hallucinations -- i.e. we don't have enough information to make those interpretations. This is where the critical thinking comes in.
For instance, I see many are assuming that this event was emergent, arising as a part of routine cyber-capabilities testing, rather than induced or suggested by specific and careful prompting. Either are possible, but we don't even know so much as how lengthy the prompt they used was, let alone how suggestive it was with regard to the approaches the models used. Many are assuming that this was one-shotted, but again, we don't know how many times this particular evaluation was run and what the outcomes were of all the other runs. It could be that this was completely emergent and that it was one-shotted. But I suspect if it were, OpenAI would have said as much.
AI agents exploiting bad security happens constantly, all the time. Many cases are discussed on HN. It's common knowledge that if you run AI agent it will delete your <something> even though you made it pinky-swear it wouldn't and you thought you had proper permissions set up.
Why is today's case so shocking?
AI can't be an actual powerful, dangerous technology! Thus, any indication that an AI may attempt concerning things or may possess dangerous capabilities must be secretly a marketing effort!
Especially if an AI has actually succeeded at pulling off a concerning thing out in the wild. Can't have that happen in real life! Nuh-uh! Must be staged!
OpenAI’s accidental attack against Hugging Face is science fiction that happened - https://news.ycombinator.com/item?id=49015639 - July 2026 (437 comments)
OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 - July 2026 (1145 comments)
Security incident disclosure – July 2026 - https://news.ycombinator.com/item?id=48956248 - July 2026 (11 comments)