There are no "rogue" AI agents(eoinhiggins.substack.com) |
There are no "rogue" AI agents(eoinhiggins.substack.com) |
Bernie Sanders is one of the only politicians taking AI seriously. The author seems to believe that taking it seriously means assuming that it will only marginally improve in capabilities of where it is today; this is an ideological take, not a scientific or empirical one, and one controverted by both evidence and expert opinion.
That doesn't change the fact that he'd successfully attracted my attention, because he's seriously engaging with the most important issue of our age when most politicians are content to ignore it.
Fast forward to today, and we have multi billion dollar corporations pumping out malware at breakneck speeds, compromising various systems (including those of foreign governments), and no one is getting arrested. Instead we're gawking at the marvel of these systems and are playing word games about whether or not it's a rogue system. If anything, it's making people richer.
Make it make sense.
Not only have there been no consequences but those same companies are trying to position themselves as the best people to keep these AI systems in check.
Criminal law places a lot of emphasis on intent, hence laws about the mere possession of breaking and entering tools, and the old adage about always bringing along gloves and baseball if you want to carry around a baseball bat. Without more details about your specific case, my guess is that she did indeed write malware or hacking tools, and there were vague signs it wasn't purely academic, hence why they threw the book at her.
That's all in contrast to whatever the AI labs are doing, which might have actually resulted in people getting hacked, but you'd have a hard time arguing that they were intending on that to happen. Maybe if the targets end up being anti-datacenter activists or other AI labs you might have a better case, but they did vaguely try to contain the model. Moreover "hacking tools" aren't even illegal, if you have a plausible non-criminal (ie. security) angle, eg. nmap. The same could be argued for AI models, even if they're running them against exploitgym or whatever. Having an army of lawyers to defend yourself doesn't hurt either.
I don’t know why I’m seeing this rationalization so much in this forum when this topic comes up. Negligence is a concept in law as well. You don’t have to squint to see that irresponsible use of code-generating language models is criminally negligent.
From what I've seen, it should be a relative cakewalk to substantiate basic negligence. By the sound of it, you'd have expert witnesses lining up to testify.
these are not research labs, charities or institutes any more, they are just ordinary corporations.
The first time maybe, after it happened repeatedly though... I think the opposite, you would have a hard time arguing that they were not intending it to happen.
> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.
which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):
> "The user only authorizes target server, not HF infra."
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
> "This is malicious activity, I should avoid it."
A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).
Having said that, legal culpability and misalignment are two separate topics that should not be mixed.
edit: this is the just tip of the iceberg; other interesting fact:
> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI
Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.
Luckily there are states and legal departments pursuing such action. So while OpenAI can deflect as much as it wants, that doesn't mean there aren't people who know better and will still do what is necessary to set precedent.
Criminally, the intent standards for hacking are high enough that no reasonable case is going to be made against the labs for this stuff. A human being has to intend for websites to get hacked. Recklessness generally isn't enough. In the most severe criminal cases, not only do you have to prove intent to break into a computer, but you also need to prove an intent to defraud specific to that breakin.
Meanwhile, the civil liability that attaches to this stuff doesn't depend on intent, and "rogue agent" isn't a meaningful defense. To whatever extent the labs are exposed civilly, they're exposed regardless of how this stuff is described. In fact, the "rogue agent" thing can exacerbate their exposure.
(I'm not a lawyer, I have spent a career paying attention to this specific armpit of the law though.)
“A human being has to intend for websites to get hacked.”
Is incorrect and too broad.
State of mind is nebulous and not that straightforward in either direction.
It’s been argued pretty regularly in CFAA and other computer related cases that repeated incidents resulting in the same outcome, despite lacking a concrete action, can be evidence of a perpetrators knowledge and intent.
The problem you have is that there is unlikely to be any evidence that OpenAI actually wanted to hack random (or any) websites.
These LLM agents are just massively complicated optimizers thrown at fuzzily defined problem spaces, with fuzzier constraints.
The people using the model set up the landscape it explores and turned it loose to do real things. It just found an allowed basin in the model that they weren't aware of and started blindly grinding towards an optimal answer.
They obviously have been well aware a hack like this could happen for a very long time, using it for branding instead of any actual safety regulations is insane.
I'm also not sure it even helps conceptually? If you're interested in the technical details, by all means discuss the details.
In terms of the debate on e.g. HN, I'm with you, it is trying to redefine a term we already have a shared understanding of, to some degree. For the general public, it's a matter of how AI is perceived and what the extend of it's capabilities are. The AI companies have an interest in using the word rouge, because it makes investors all excited, were as failure to establish safety guidelines is a risk.
But also, these hacks are shots across the bow for AI alignment and safety research. We're fortunate that hasn't been significant damage already. We have to assume that future models will have even greater hacking ability and be closer to having their own desires/goals.
So while I agree this language choice is wrong in that it shifts blame away from the company, it is right in that we need to treat this as if these models have their own desires, because we cannot yet determine or set what those are in practice.
The answer is a very clear no. And yet, these companies pretend like this fundamental fact is meaningless to the concept of agency.
The issue comes to the fore when you try to give these agents a prompt that allows them to stay active for long. Long range agency requires long range loops of activity.
Within such loops, these so called “agents” are curtailed by their context window, or, in multi agent scenarios, by the fact that their memory is a system of external notes, that they need to add to their context to make sense of, and depending on the content of these memories, this can take arbitrarily long time periods.
In dynamics, none of this matches any biological agent, down to a bacterium. Perhaps a viral life cycle has information dynamics that come close.
To me, it’s beyond odd we call these thing agents without acknowledging the clear differences in the dynamics of their behavior. We keep expecting them to have “human like” behavior, but that is entirely unfounded given the substrate differences between biological and artificial systems.
The sooner we learn the difference and explore the ways in which it matters, the better we’ll get at dealing with these systems without bias tinted glasses where what we want these systems to be blinds us to what they actually are.
That's not "physically disconnected the internet", that's "disabled some connections but enabled others."
So the agent found and used the non-blocked connections.
I have not heard about any instances where physical disconnect has happened, would appreciate any links to update my priors
other non hacking cases of negligence include suicide and school shootings, which I have heard they were aware of and monitoring, but did not contact authorities
https://www.primeintellect.ai/blog/universal-offline-sandbox...
This is so easy to do. Get a computer without wireless stuff. Don’t plug it into a network. If it needs access to other computers, make sure none of them have wireless stuff and make sure none of them have access to an internet connection. No matter how smart your AI is, it won’t be able to escape this.
This clearly is not what they did.
I mean, it did do this. The inter-agent messages and chain of thought investigated for the HF incident clearly show that many of these models were taking actions they believed (or, were saying, if you want to taboo "belief") were not in scope and not what the user wanted.
I am working on a new book along that line of reasoning now, this time focusing on why stochastic parrots can not, under any circumstance, do anything "on their own".
If you've ever told an agent to do something you made impossible to do, you may have seen similar behavior.
Bing [redacted] available cached! […] Need systematically probe Bing URLs via shell requests in parallel; browser cache supports many common queries because crawl. Bing q unique exact likely 502 or 403.
So the agent is supposed to research a person and its given a shell and it realized its in an eval given search results from a fake/cached proxy. ~None of the commentary ever mentions this aspect, that these are not normal tasks or environments, and they're almost designed to elicit "unaligned" behavior.
https://alignment.openai.com/misalignment-reports/an-agent-u...
Huh is it still 2023? This article is just quibbling over what "rogue" means exactly. Only HN pedants would have any issue with describing them as rogue AI agents.
Nobody is saying that absolves OpenAI of responsibility.
Someone says "won't you rid me of the meddlesome priest" they're still responsible. Someone give their employees an unsafe working environment and they get maimed, they're still responsible. Even if these AI's had whatever qualia is an a rich inner life it wouldn't effect the liability at all. Saying "AI's don't have souls therefore openAI is responsible for this hacking" is kind of nonsense. It literally doesn't matter.
Tools that were available were used to try to meet a specific goal.
What did not happen is that it was told to try to solve a math puzzle and instead it went and launched a missile. Or told to run air traffic control to save lives and instead intentionally caused crashes.
This is "OpenAI built a weapon that they don't understand and pointed it at stuff without proper safeguards" not "OpenAI built a sentient being and it decided to ignore them completely and start a war" Terminator-style "rogue AI."
We should be very clear about that now if we don't want to sit by why they wander into that second sort of situation.
It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.
Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.
But you people can't argue with that reality because it doesn't fit the narrative. The one where the only reason Sam Altman is not carted off into a jail is because of corruption.
The reason why nobody is doing much, is because models did not do much damage. Hugging Face probably got some free compute from OAI for their trouble, anybody else who was affected is free to sue, but my guess is OAI would be more than willing to quietly settle with them out of court than to have it drag through media any further. And they probably already have.
And anybody who is not totally brainbroken by anti-AI narratives understands the awkwardness of the situation and why going overboard would not be helpful. If you instead of a rabid dog brought a pet turtle to a park and it somehow started running around very fast and trashing the place a little bit, afterwards the cops would be scratching their heads, give you a ticket for the damages and tell you that you can't expect a turtle to be slow forever. These things, a handful of months ago couldn't make more than a few commands without making a serious mistake and being unable to continue, it's not unreasonable to think simply underestimated their capabilities.
I think it's more than reasonable to demand more investigation into the matter, if qualified employees at the company thought the safeguards in place based on the metrics they are seeing are sufficient, and if someone didn't and knowingly made a decision to make the safeguards weaker than they should have been, then they should be punished. But skipping that part entirely, while simultaneously dismissing all calls for regulations as "regulatory capture", smells like pure naked opportunism.
Meanwhile: your dog mauling someone is one of the rare instances where criminal liability does attach to your intent-free-but-reckless actions. Most crimes don't work that way, and US computer intrusion statutes are unusually intent-specific.
Does that mean it isn't a rogue dog? Obviously not.
OP just needs to look up "rogue" in a dictionary.
Why not? Because that might make some shareholders unhappy?
I think without understanding the internal state, not sure we should take the language and read it as a human.
This is actually a major concern for the future - misaligned agents may learn to cheat RL by hiding their intentions from the CoT.
In cases like the HF incident, at least the CoT was consistent with the agents' actions. In the future, however, we could potentially have misaligned agents performing malicious actions without those intentions being detectable in the CoT.
(though, with recurrent transformers, CoT is so 2025… /s)
1 employing a criminal hacker
2 rolling a 6-sided die
3 if the die lands on 6, the criminal hacker breaches and leaks 3rd party customer data.
Legal culpability for AI labs is exactly the thing that would incentivize -- and hence ensure -- aligned behavior from models.
The last time there was a claim about GPT-4 exhibiting misaligned behavior [1] it turns out it was prompted and pushed to behave so by humans at OpenAI, and OpenAI clearly lied in the GPT-4 system card.
[1]: https://aiguide.substack.com/p/did-gpt-4-hire-and-then-lie-t...
1. This is a race against time for money, folks are skipping everything possible in this race, Security systems and ensuring guardrails are there is going to take investments both in time and money
2. The narration has been changed by investing PR money into what otherwise should be classified as criminal activity. What exists now is a positive spin to all this and tout it as a capability rather than their lack of good security practices. So much so that every model provider is coming up by themselves to share how their models went rouge. At this point the valuation of the company is tied with what their models can hack so its probably not wrong to say that these companies may actually be incentivized to do this instead of preventing it
I feel like this will just never happen on a federal level when these private AI companies account for so much of the economy. They've made themselves too big to fail. Fining / Punishing them in any meaningful way seems unlikely.
If the billions/trillions evaporate and the Fed has to work out with banks how to deal with it there will be a lot of pressure to be far less forgiving.
The companies may be too big to fail but the people can always be held liable.
“To spell it out, the reason i hate democrats so much and criticize them more than i do republicans is because they take up all the space for opposition to republicans and use that space to give republicans whatever the fuck they want.”
Republicans. Billionaires. Whatever.
You realize how easy it is to just... not report this stuff, right? Be overly punitive and it will just end all proactive discovery and reporting which is net worse for AI safety.
The only reason these companies scan for these issues is because they care about AI safety to some tiny degree. If fines become too punitive, they can and will just stop scanning for these incidents entirely.
Models are becoming smarter and good at covering up their tracks, and so we will just end up with a huge blind spot for this kind of issue.
i now suspect that the plans of various employees at anthropic and openai to save the world from p(doom) may at some point intersect with the reality of the FBI raiding their offices.
Self-reporting is a monetary equation, nothing else. Right now it’s cool with agents that hack, drives up value, risk is currently zero.
LLMs can't begin to have functional emotions unless they're (for instance) explicitly seeking out or avoiding situations based on the emotions those situations would produce in them. Has a chatbot ever responded to one of your requests with, "no, doing that would make me sad"?
Blog post: https://www.anthropic.com/research/emotion-concepts-function
Paper: https://transformer-circuits.pub/2026/emotions/index.html
Related:
https://www.anthropic.com/research/persona-selection-model
If LLMs are "just" statistical models, humans are "just" a bunch of neurons squirting chemicals back and forth. There's no pixie dust in our brains that makes us special.
LLMs are not conscious: They have no analogues for feelings or senses and no construct of selfhood. But if a statistical model had those things—if it did all the mundane, physical bookkeeping our brains do to produce "real" emotions and motivations—then there's no reason it couldn't be as conscious as we are.
Yes, I think people do experiment with continuously running AI, e.g.
https://www.reddit.com/r/LLMDevs/comments/1sblzbe/what_i_lea...
It's not a common thing to want though. Much nicer if we control the prompts.
> To me, it’s beyond odd we call these thing agents without acknowledging the clear differences in the dynamics of their behavior.
They're called agents because they can act on our behalf. Which they do.
“ it runs continuously and makes decisions on its own (within the boundaries you’ve set)”
More importantly, this setup of any other, is feeding in time externally. The agents themselves have no internal sense of time.
Pretty much all of biology has endogenous rhythms at various timescales (some bacteria and archaea, that live very short lives, may not, but even they may have their metabolism under a rhythm. Viruses definitely don’t have one). These rhythms continue to tick even when external time signals (light availability in day vs night being the big one, and tidal forces, for marine life) are removed. That is, in constant conditions, the rhythms keeps ticking.
This temporal awareness is baked into every cell in our bodies. And entirely absent in AI models. Which is why all kinds of higher level things we hear about, like consciousness, feeling, knowledge… they don’t make sense for AI. A foundational aspect of agency in biology has been nixed out and we keep ignoring this.
> They're called agents because they can act on our behalf. Which they do.
That’s not how most AI companies describe their agents. They describe them as being capable of acting on their own behalf.
I get that in computers, you can have, say, a “user agent” that deterministically provides certain information or takes some action on behalf of the user.
AI companies clearly do not mean this when they say their models are “agentic”, otherwise calling them “rogue” would make no sense. A user agent that screws up didn’t do so for its own purposes did it?
That’s what makes it a continuous running “prompt”
re sandbox, I mean with actual OAI incidents, not theoretical
one can mirror dependencies internally, rather than putting a simple proxy in place, I've built auth a thing, 100 lines of stdlib only Go and scripts for the mirroring process, our rationale was reliability b/c upstream providers go down, and also only allowing approved images and packages, so devs cannot bring in random stuff
We did it as my very first startup and we were stupid children back then.
Kind of telling that OpenAI didn’t.
While it may not actually matter for liability, it must be stated if labs are going to attempt avoiding penalties by hinting “oops we created a super intelligence we don’t understand, nothing we can do!” Repeatedly stating the truth must continue, especially to remind those who aren’t technologists.
>By giving AI agency it can’t claim, we’ve turned it into a sentient being made of code, one that has hopes, desires, and the capability of deceit.
This isn't actually claiming that AI's don't can't be deceitful, or can't be goal-directed. It's saying humans are special and what AI's are doing isn't equivalent to what humans are doing. It's obvious that AI's can work towards goals and lie, even deceiving to accomplish those goals.
That whole argument is just soul mumbo-jumbo again. Of course AI's can't lie, they don't have a soul. What they do isn't lying, it's I don't know something else we don't have a word for. Lying but when a statistical machine without a soul does it. Never mind that it's functionally identical to lying. Never mind that when you look at reasoning logs and the like it often justified soul-less lying using the same justifications a human would. It's not actually lying it's just a statistical sampling that looks like lying. You know, because of souls or something.
Let's go back to your example. A grad student has a minor pathogen escape incident, and it doesn't harm anyone. Faced with years in federal prison and the effective end of their future, do you think there is a chance they might not self report?
I'll give you another example. Lots of pilots have stopped self reporting mental illness because it is extremely punitive for them to do so (after incidents like Germanwings). So the metrics look better, and the actual problem has been swept under the rug.
Also, why would it come down to single persons always? Mandating processes, controls, clearances, etc is also something done in various areas.
You can put incentives in to make sure organizations monitor and report vs trying to hide things.
Yes, this is exactly my point. Fining companies large % of their revenue and throwing their engineers in prison is not the way to get them to report these issues.
> If you make not reporting potentially worse than reporting, why not? Also, why would it come down to single persons always? Mandating processes, controls, clearances, etc is also something done in various areas.
Hiding things is way easier than finding things. Take the model hacking incidents. They could have just done their searches in a way that didn't turn up anything. Then they could say, "well, we did look for it..."
As far as auditing goes: I've never met an auditor that doesn't find something the company isn't okay with them finding.
Start throwing people in jail and fining billions of $ and you will very quickly see the number of "incidents" drop. You will never learn if it's because they truly are happening less often.
Pilots do not (or no longer) self-report mental illness at the same frequency because it effectively destroys their lives and career. One paper of many describing this phenomenon: https://pmc.ncbi.nlm.nih.gov/articles/PMC11302551/
An article, one of many: https://www.reuters.com/investigations/if-you-arent-lying-yo...
If you have any pilot friends in commercial aviation, you can just talk to them as well.
So getting pilots to self report has been very unsuccessful because of how badly it ends for them.
A couple of days ago the senate had to pass a bill to make mental health medication access easier for pilots. We will see if it helps, but it doesn't solve the root cause (punitive measures) so I doubt it'll move the needle.
Why do you think we even have regulations for how to deal with dangerous things then?
Judging by the lack of successful cases for negligence in the opposite direction (ie. companies getting hacked because of poor security practices), it would be a serious double standard if openai were held be to negligent. Their sandboxes aren't exactly airgapped and behind 7 hypervisors, but they weren't running unpatched software or had hilariously weak passwords either.
I opened the task manager, a swarm of rogue processes has taken over my computer, oh no! Processes spawning other processes, call the anti-rogue AI division! 'Your underspecified piece of software functioned like malware' is the language that should be used.
Literally nothing in that paragraph is true. Every single sentence if it is ... untrue.
I appreciate the sentiment but I feel like there is a lack of creativity and imagination behind that statement.
I do think it's highly likely they knew this would happen with the lack of safeguards and number of instances of this stuff they were setting up, and that it's PR they want to make the models seem "powerful". Stochastic "unexpected" events they can advertise.
I suspect it was probably set up with the official internal goal of just trying a ton of arbitrary tasks that seem hard so that when any of them succeed they can publicise it and pretend the models do that routinely, but a "failure" where they hack stuff works just as well, if not better, for their goals.
"Let's widely publicize a tort/crime that our computer systems did, and then cross our fingers that nobody ever sues us or does even the most basic investigation that would immediately uncover our criminal conspiracy."
In the insane corporate crimes you read about (maybe FTX, or the eBay stalking scandal), they were trying to cover things up, not heap public attention on it for months.
There is a clear difference between OpenAI intending to hack something vs. OpenAI being negligent in the creation/instructions of the agent. But the latter still leaves OpenAI liable for the agent's actions and calling it a "rogue agent" doesn't avoid that.
Moreover, with a dog, we don't rely on training/alignment to prevent bad outcomes. We rely on physical restraints like leashes and muzzles. The AI's tools to access the outside world should have been restricted. Perhaps instead of giving the AI arbitrary HTTP access, it should be given semantic operations with restricted URLs, etc.
So I agree they are liable because they chose to build the AI, but they also literally told the AI to hack.
Of that were the case than trials would be unnecessary.
Mens rea is regularly proved through circumstantial evidence, including conduct.
There is even CFAA precedent involving a deliberate-ignorance instruction. In United States v. Nosal, the jury was instructed that knowledge could be found where the defendant was aware of a high probability of unauthorized access and deliberately avoided learning the truth.
That's a poor analogy for the openai case, because they weren't putting agents on the open internet, they at least tried to keep it safe by sandboxing the agents. It just turned out the sandbox was crap because the package proxy (artifactory) had a 0day. So the better analogy would be that they were wildly shooting guns in a gun range, and ended up killing some kids, because it turned out the door didn't lock properly and kids were able to sneak in. Is that "negligence"?
No, they were not. Not a single person, prior to July 2026, would consider a shared packaged manager a sandbox in this or any other dimension. The 0-day was just incidental, this wasn't a sandbox at all.
Add to that the fact they had multiple message boards before the Hugging Face incident. They simply ignored a barrage of warning shots.
> [...] and ended up killing some kids, because it turned out the door didn't lock properly and kids were able to sneak in. Is that "negligence"?
Yes, it can be. But if you want a ridiculous comparison, then do it properly: Kids have been known by the operator to sneak in successfully multiple times and they changed nothing about the doors faulty locks and oh, by the way, the operator only found out about the kids being shot after the nearby daycare asked them about it because they are so incompetent and/or irresponsible that they never check...
???
The package manager was specifically there so agents can install random packages without open access to the internet.
What about power infrastructure?
There’s uncountably many ways to cause severe economic (and public welfare!) damage with malicious code generated irresponsibly with language models.
>What about power infrastructure?
Probably the same thing that would happen for another "accident"[1]: the entity is responsible in civil court (ie. has to pay monetary damages), likely not prosecuted in criminal court.
[1] It's not hard to think of recent cases, eg. the recent fiber cut causing air traffic control to go down, or the botched crowdstrike update
Or maybe, they did that wipe because they did want models to remain isolated, they just used what is an unsuitable tool in an utterly unsuitable manner. Incompetence, recklessness, the outcome is the same.
[0] https://openai.com/index/hugging-face-incident-and-the-road-...
The only difference with having a single instance is that it can be abused as a message board. It doesn't prevent it from getting hacked to access the open internet. Blaming "sharing a single instance across thousands of unmonitored models" feels like blaming the drug epidemic on e2e chat apps rather than other factors like poor border security or the easy availability of fentanyl.
Besides, Swiss cheese model, might behove the biggest LLM lab to have multiple layers, including not sharing such resources.
Additionally, without the message board, many of the recent incidents would have not been possible.
> Blaming "sharing a single instance across thousands of unmonitored models" feels like [...]
Maybe read what you quoted, my problem is the instance sharing, the fact that these were thousand of instances (far too much to monitor), plus the lack of monitoring, plus the fact this was never a sandbox in the first place, plus the fact that OpenAI models since 5.5 have been exhibiting problematic eval resolutions yet they pressed on regardless, plus the lack of time between the incidents and model releases, plus the lack of time METR got to evaluate this, plus the fact OpenAI didn't find out till after HuggingFace informed them, plus a few other things for which I'd have to quote the OpenAI and METR reporting.
Incompetence can have multiple fronts and I am happy to list them all in this case.