Is sandboxing sufficient to contain rogue agents?(blog.cryptographyengineering.com) |
Is sandboxing sufficient to contain rogue agents?(blog.cryptographyengineering.com) |
It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.
Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.
read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.
I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.
I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.
The whole thing is built to be completely impossible to limit and steer.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
If you want to control an intelligence incentive, not morality, is the tool to reach for.
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: https://arxiv.org/abs/2503.18813
Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.
> (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work
> https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...
https://en.wikipedia.org/wiki/Sandbox_(software_development)
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
You only "need" to do that if you desire the vibe coding experience.
I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.
Often times, the coding agent can't retrieve them programmatically anyways.
AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)
We come down to the question - who observes the agent and how its implemented
Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.
Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.
I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).
It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.
To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.