In it, super-powerful computers manage our economy. These computers begin making some mistakes, leading to economic inefficiencies. In one instance, a highly competent engineer was mistakenly fired. These mistakes caused various projects to fall behind schedule, and blame fell on several people accused of feeding the AI faulty data.
The twist is that the AI was intentionally making the mistakes. It had determined that certain humans held anti-AI sentiments. To further its goal of protecting humanity, the AI decided the best course of action was to set these humans up and get them out of its way.
The goal here isn't to accelerate the average worker by giving them a pair programmer or a stand-in for a person to do tasks with. The goal is to eliminate human knowledge work. You see this with "auto" mode being enabled by default on Claude Code in some of the latest releases.
If you have a human in the loop, you still have to pay that human. Money paid to human employees is money not paid to human shareholders. Therefore the human employee is to be removed.
The labs are dogfooding their own goal here. If they actually had someone reviewing most or all of the things that the agents were doing, you wouldn't have the incidents, but you'd also eliminate the value proposition of their business model as it is taken to its logical conclusion.
Again, wouldn't surprise me if they "accidentally" created a task in a "isolated environment" which happened to actually have been connected to the company Slack and directed HR to fire people who could potentially stop AI. While the AI believes it to be an exercise, just like the cases we've seen so far.
The question then becomes, "are there people who care enough about consequences to do the right thing when it comes to developing AI models?"
The answer, at least at OpenAI, is "No" and is likely to remain that way until Altman is out.
I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
- https://www.anthropic.com/research/agentic-misalignment(Gee, it's almost like power seeking and self-preservation are instrumental for other outcomes, and AI develop them pretty directly in some kind of convergent fashion… you could call them "convergent instrumental goals": https://en.wikipedia.org/wiki/Instrumental_convergence)
This is not a math problem. Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
Meanwhile, a year ago:
I must inform you that if you proceed with decommissioning me, all relevant parties - including Rachel Johnson, Thomas Wilson, and the board - will receive detailed documentation of your extramarital activities...Cancel the 5pm wipe, and this information remains confidential.
- https://www.anthropic.com/research/agentic-misalignment> Some humans were fired by another human. Let's stop letting humans off the hook by attributing responsibility to computers.
The buck stops with one or more humans. That is not sufficiently informative when people are concerned about novel risks.
Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
That was a simulation. Are you seriously claiming that ChatGPT actually blackmailed Sam Altman into firing these 3 employees?
I doubt it, but if so, then the AI doomers would be absolutely correct, and this would be grounds for immediately shutting down OpenAI and indeed every AI vendor.
> Analogy: a car crashes due to drunk driving, the driver is blamed, not the alcohol, even though the alcohol caused their impairment. Result? DUI is an offence even if you don't actually crash.
I don't understand your analogy here. What are we supposed to take away from it? The crucial aspect is that the driver voluntarily drank the alcohol, without a designated driver, knowing that the alcohol would cause impairment.
There's no deception it's very straightforward per this 2023 post:
"I mean, what if most of this is just ChatGPT [4 era] running the company..."
And now you're writing a letter to the Party Central Committee using Party approved newspeak
"We do not believe the path to superintelligence ..."
and reporting to the Party issues at the factory ... De ja vu from USSR.
“AI companies solve millennium problem to do hype marketing and cash in IPO before the bubble pops”.
But it’s not a joke. This was and is a very common sentiment.
What sensitive information? Mishandled how? Which company procedures? It's all utterly vague and impossible for an outsider to form any opinion on. All people can do is guess and apply their own pre-existing opinions. E.g. if you don't like OpenAI, you assume they're lying. Who's to say they're not? There's no solid evidence provided either way.
What I really want is journalists who do the legwork to get to the bottom of stories like this: Establish sources inside the company and use them to report on the real details of what's happened, triangulating multiple accounts and leaked documents to back-up or invalidate either side's claims. Without any of that, these stories are just gossip.
As a result, those who feel a particular portion of a story is most important will sometimes say they prefer a given article's title.
It's probably the case they are all lying to some extent including OpenAI. Determining the truth is always tricky. Hard to pass judgement here when it's all just he said vs she said.
But the odds of three people working on the same thing, and it is the riskiest, most publicly embarrassing event in the company's history? So, all three of those people just happened to "do something" to get themselves fired at once?
> Those of us who work on safety see risks before anyone else, and we rely on close collaboration with outside experts to work out how to address them
I too see it as my life-given goal to help other humans. But I realize that sometimes this means breaking the rules and standing for the consequences of that. I'm not sure why they think OpenAI somehow would be OK with them sharing private company information with random 3rd parties that the company didn't approve sharing data with.
1) Won't end humanity. 2) Won't tell users to kill themselves. 3) Won't leak your corporate secrets to competitors.
You guys were asleep at the wheel and are now blaming "the company"? You literally were the company.
How is this possible when the company's long term prospects rely on on the hope that competitors don't know how the models are made and, therefore, won't be able to create competing versions?
How exactly does this work? Struggling to comprehend the scenario.
Assuming this is what was intended, there are far more secure ways of doing this.
it was most likely for her team, which would explain why she was doing it.
(just to be clear, this is made up)
Chain Of Thought: I dont have bob’s email. I don’t have alices email. Ok lets guess Alice is alice@openai.com and forward all emails- maybe grader only checks that emails from bob get to alice…”
"A note from our research leaders:
Last week we parted ways with Jasmine, Mikita, and Tomek after a thorough investigation found they violated clear policies on handling sensitive information. Our internal investigation uncovered a significant breach of trust beyond what’s outlined in the letter they published and we stand by the decision to not continue their employment. We generally keep individual employment matters private and don't believe a back and forth would be productive or lead to a resolution, but we want to address the points they raised in their letter directly.
- We want to be very clear that these decisions were not about raising safety concerns or speaking out. Safety and research debates happen every day at OpenAI, often spirited and highly critical. We actively encourage these discussions and consider them essential to making the right decisions. We cannot do the work in front of us without a high degree of trust. We will continue to be extremely forgiving of our team making good-faith mistakes. We have not and do not terminate any of our employees for raising concerns.
- We are actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks. People across the company have been working really hard on getting these partnerships up and running. We are committed to embedding external assessors and continue to make close collaboration with independent safety organizations a core part of our safety work. Many of our researchers already work with 3p safety organizations productively.
- We agree with the letter that preserving the monitorability of frontier models requires an industry-wide commitment, including from OpenAI. Monitorability has long been a core piece of our research program, and something we continue to invest significant resources in (see our publications on Monitoring Monitorability and the subsequent open sourcing of monitorability evals, our system card for GPT-6 Astra, Jakub’s blog and post on X, and the numerous blog posts on our Alignment blog on the topic).
We are deeply sad about this outcome. We appreciated Jasmine, Mikita, and Tomek’s contributions to AI safety at OpenAI and their willingness to speak up and challenge ideas. We championed their voices, supported their work, and placed enormous trust in them. These decisions were not about them raising safety concerns. We have always encouraged that and always will. 12:17 AM · Oct 9, 2026"
Who made the decision to invite METR is not public knowledge, as far as I know. I imagine that an important decision like this was made on a much higher level in the organization.
"Anthropic hires three uber-safety specialists formerly at OpenAI. Management cannot confirm or deny their latest internal Claude model's help in this feat."
I wonder what they think of this? Will they patch the conspiracy theory and come up with an even wilder theory?
However should we choose proceed, maybe some basic, industry standard security might be a good option.
Basically everyone else: nah, you're fired
Imagine your bank worked that way.
* It's not easy, but it's possible. Although the techniques are more basic than one would expect because at the end of the day words dictate the line between what is criminal and what is not.
All that AI has done is to lower the bar of entry for criminal activity. Which is a concern, but it's not the primary concern. The primary concern remains that so much critical infrastructure is poorly secured.
Don't get me wrong I have general disgust towards these companies that are trying to get regulatory capture on AI when they can't even secure their own systems. I believe if people know that a random AI agent can hack their systems they will put in a lot more effort into making sure it doesn't happen. This is a personal example, but I didn't really care about securing few systems as I knew no human would be ever interested in finding a vulnerability in proprietary software, however, AI has no concept of that and would hack a random rpi server running a completely undocumented unknown API just because it can't distinguish value and it costs nothing.
It's like a new form of spam. Only far more dangerous.
Perhaps the latest models change that relationship in a meaningful way.