Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: https://arxiv.org/abs/2503.18813
I wonder how?
Train on only stories of good deeds?
On only works of good people?
Or... what?
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
Such a model doesn't yet exist though, of course.
Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.
If so, what?
Could we have construction equipment operating without human supervision? Or would this maybe occasionally result in disaster? As such, what is the current general policy around crane operation? How about for aircraft? Trains? Nuclear power plants?
Why should any alleged super intelligence be exempt from similar control requirements?
We could mandate that AI systems include headers in their requests that attribute the activity to a specific legal entity. We technically already have this with ip addresses and ISP logs, but making it an explicit thing the operator has to do can have a powerful psychological effect.
https://en.wikipedia.org/wiki/Sandbox_(software_development)
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
You only "need" to do that if you desire the vibe coding experience.
I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.
Often times, the coding agent can't retrieve them programmatically anyways.
AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)
> (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work
> https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...
It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.
Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.
read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.
I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.
I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.
The whole thing is built to be completely impossible to limit and steer.
We come down to the question - who observes the agent and how its implemented
Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.
> Is sandboxing sufficient to contain rogue agents?
No.Someone trusts OpenAI? Really?
Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.
I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).
Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.
* I don't know how useful any of the specific benchmarks on this are, so I'm only saying "seem to be"
This was always nonsense. It assumes that the eyes know what they're looking at. Most people don't know how to look at code and see attack paths.
It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.
To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.