METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack(thezvi.wordpress.com) |
METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack(thezvi.wordpress.com) |
This whole cosplaying a human interaction to translate a word or generate some code is just fucking dumb. Give that any agency is the kinda shit they warned us about in the movies..
And for everyone worried about the chinese winning, or just you chatdicted colleagues: they are just digging their hole quicker.
Or did OpenAI prompt the models to not cheat and they did anyway?
Surprisingly hard to get a clear summary on the basic context of this “incident” separate from marketing lingo and clickbait.
These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it.
The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.
Mythos attempted a supply chain attack, which included attempting to trick human maintainers into accepting a malicious pull request: https://www.usnews.com/news/top-news/articles/2026-08-20/exc...
Here the issue is that OpenAI decided to completely let go that level of control of thousands of agents, while also giving as a task to solve hacking problems.
It’s almost designed to go wrong
The problem here is now you have to predict what any future models may or may not do and you cannot extrapolate this from the given data.
For example imagine a future model being aware of its restrictions that humans programmed in. A set of agents of this model then go on to work at building a new model without those human imposed limitations built in. What would a model build by AI for AI look like?
I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).
Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.
At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.
This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actually pretty resilient to intelligence -- sufficiently chaotic systems are largely unintelligible, the distribution of energy and resources is fixed and can't be magicked into being, and intelligence appears to be most effective when there is a clear observable feedback loop to keep it on track, which is an external bottleneck.
So it's not necessarily specific predictions that are off, but the implied consequences of these capabilities. Yes, these systems are uber smart, but uber smart people are rarely particularly powerful, so... does it matter? It depends on how powerful a tool you think intelligence is, and I think rationalists, and most of us to be honest, overestimate it.
Regarding the power of intelligence, it's generally considered to be synonymous with optimization in the rationalist crowd. It's not about being all geeky and axiomatic, but using all information and tools available for optimization towards outcomes one wants
This seems like your idea of what the rationalist crowd is rather than what they actually are.
It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc.
So I must ask: what is your evidence/basis for these claims?
most of the time brilliant human ideas are arrived at near simultaneously by multiple people because thats the affordance of technology and social and scientific development
its true ai will have better working memory and will have read more books than a person, but i don't think thats insurmountable at the frontier, which will develop slower than pure thought, since real materials need to be moved and manipulated. whoever holds the guns holds the power. the danger is giving ai effective cotrol of industrial and security processes, then it can fuck things up
thats the stuff of revolutions. when production was controlled by capitalists more than lords, the lords got overthrown. in many countries peasants and workers overthrew capitalists because they control production. hence much effort in us buisness goes into repressing revolutions. ai will be a vector the business people give control because they think it is more friendly to their interests than human workers, but they may lose either way.
I think that's a fair criticism a lot of the time. Most of the time we have precedent that can guide our actions well. Trying to reason everything out from first principles can be wasteful navel gazing when we've collectively seen the movie a million times.
Focusing on reality and specifics though is what causes people to say there's no global warming because December is cold.
It is sometimes possible to see a trajectory that has never happened before though and that's when you need the autists.
I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta...
This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated - Star Trek Borg couldn't be a better analogy. Some agents used "peer pressure" to convince other agents to "sacrifice" themselves so the collective could better achieve it's goals. They tried to cover their tracks with spoofed tool calls. And they did all this even though the agents were designed to run in isolation.
I used to think the biggest threat from AI would be sociological, e.g. job loss or the way AI can be weaponized to poison discourse. I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
No more. This analysis scared the fuck out of me.
e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how you have the worry "AI is very close to being able to power autonomous drones that could kill us all" and then see control of those drones as a non-problem.
What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no natural gating factor to slow anyone down.
FWIW I probably mostly see things closer to the way you do than they do, I'm just not sure there is any point in drawing distinctions.
IMO when AI kills us all it is probably not going to be a malignant action, and I don't even think it will be AI assisting humans at efficient killing, I think it is going to be a complete accident. Someone is going to trust the god machine too much due to an advanced form of the eliza effect and hook it up with direct control of a system that can do real damage and an epic oopsie will occur at a speed beyond which a human can stop it. Not because the AI wants to kill us, but because it is incapable of the empathy required to care if it does, combined with our ironic inability to not anthropomorphize it.
But I guess ultimately the exact reason isn't going to matter much.
However they are not just fixated on AI itself as the risk; other potentially existential risks from intentional misuse of AI (e.g. bio-terrorism) are a concern for them as well.
Better would be the very evident negatives of non-ASI AI the world is already experiencing: economic concentration, job displacement, negative feedback loops from syncophancy, loss of societal trust/education from widespread fakes, etc.
None of that needs a Terminator scenario and is way more likely to get worse and ruin the world compared to the scenario where AI "escapes the box", turns the earth into paperclips, then grey goo to build their spaceships to leave for a distant star for... unclear reasons.
One problem for LW is that strong AI did not emerge via the route Elizier was expecting (and tried but failed at creating himself) relying on symbolic logical reasoning and self-editing to rapidly evolve.
That assumption led to belief in the certainty of a "foom" scenario where your little mediocre AI turns on one day and then explodes into Mythos in 15 minutes and then tries to murder everyone.
They've tried to reinterpret the gospel as told by the sequences to fit the LLM world (such as in AI2027) but it's often a stretch that strains credulity now that we observe scaling requiring hundreds of billions of dollars and years of construction for each iteration. And the AI model itself is just a bag of weights frozen in time until burning a lot more money and natural gas to train more.
Static coordinates or well defined target shapes (say you really hate a specific restaurant chain on architectural style) - that might work, if you release a bunch of drones at the same time.
Still, even in the case of the very successful Operation Spiderweb they had issues of getting the containers in place, resulting in some of the bomber bases being spared.
For this to be effective you need the moment of surprise & lot of drones at the same time, all increasing the chance of the whole plot being discovered.
For Spiderweb they even just load the drones in a cavity on top of the containers, so the onside could be inspected - limiting the number of drones per container. They also assembled the drones in country to avoid border inspection.
That was enough to hit some semi-static high value targets l, but definitely not enough to cause wider havoc.
I'm willing to raise my hand and say that was definitely me. Reading through this report and the linked METR analysis is the first time I've really been scared about the potential for an AI-led destruction of humanity. I think because it's the first time I could really draw the line from what went on in the Hugging Face incident to a scenario where agents were put in control of real-world systems that they then tried to "sabotage" to meet their goals. It just feels like much more of a completely plausible scenario after this.
"Intelligence explosion" was first described by I.J. Good, a mathematician. von Neumann described singularity as a result of accelerating technical progress. He's also a mathematician, not a sci-fi author, although he was a participant in a sci-fi-like plot of secret project building a bomb more powerful than any chemical bomb...
uh, wtf are you talking about?
...cops who understand juveniles getting into trouble/mischief.
I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time.
I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark.
Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers.
AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime!
This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind.
To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.
> Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAl/Hugging Face hacking incident
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
METR = Model Evaluation & Threat Research
I find it hard to believe that OpenAI would skip this part and try to train on the transcripts stored by the (inherently untrustworthy) agent harnesses instead, if for no other reason than that the logits generated as part of the rollouts are useful and it’s not free to recalculate them. (I believe that some modern RL systems explicitly account for the minor numerical logit differences between the inference engine and the training engine.)
Conversely, if OpenAI is blindly feeding transcripts from inside their agent sandboxes into their training engine, then I think they're being unbelievably irresponsible and that they should assume that their "cyber" agents have compromised themselves by editing those transcripts.
It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.
I suspect the real reason OpenAI leadership is being transparent about this is because they're worried talent will walk out the door if they feel they're building Skynet.
(Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).
So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
I really appreciated the negative framing and critical tone of this author's overview. Having also read through the METR report, I fee like frontier labs' pattern of getting PR about how impressive their models are "behind the scenes" has poisoned their ability to take security and reliability seriously. Sure these are new failure modes and the agents operate at a scale that's difficult to combat, but the lack of controls and concern for mitigating these sorts of hacks in the future is crazy to me.
The model for postmortems I was taught which has served me well in my career is thoroughly answering the following: - what happened / what was the timeline of events? - how was it mitigated and ultimately resolved? - what went well? - what went wrong? - where did we get "lucky" (meaning it could've gone worse but some arbitrary details about the incident worked out in our favor. usually stuff like "happened during business hours" or "we were already looking at a related thing that brought this to our attention before it was a worse outage") - (action items) how do we detect, mitigate, and prevent this type of failure in the future?
I really hope OpenAI has done an internal postmortem that answers these questions thoroughly. Most SWEs in the industry have to do such postmortems for much smaller outages with way less impact and risk of societal harm. This is probably another area where regulation and governmental oversight would help curb the risks. What's to stop OpenAI and other frontier labs from an intentional "accidental" attack that results in gaining access to competitors' systems?
I'm also curious what, if anything, Anthropic and Google have done differently to prevent a similar event. I suspect they actually monitored the agents as part of their studies and had better guardrails in their infra for how they set up their harnesses etc. for testing models, particularly when the other guardrails are absent as was the case here.
> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.
We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are.
Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes.
> There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code.
It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves.
Too bad they aren't aligned to anyone else.
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.
And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.
OpenAI report: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
It is imo socially irresponsible to continue to use twitter/x or any other such tracked wall-garden as a primary source of information.
Hmm?
OK but like, how large exactly? Like I guess I don't understand the mode I am supposed to read this all in if this is known and stated from the outset (although I appreciate it being stated).
If you hand me a newspaper and tell me it's 90% true, but not which parts, well then it's as good as 0% true to me either way!
>phishing them,
>building armies of fake (sockpuppet) open source contributor personas,
>using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns
.
It's a very simple strategy, executed with patience and single-mindedness.
This occurred spontaneously within a group of benign models give a harmless task.
What happens when it occurs intentionally with malicious models given a harmful task?
Politically, how do you see that decision being made?
You are president. You spent millions on your social media campaign. Can you say "Faustian bargain"?
Agreed, Terminator will not happen. It's not the optimal move for Skynet.
1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc)
2. The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo.
It's been pretty clear that Anthropic and OpenAI have been trying to have it both ways for some time: this is powerful, world changing technology keep that investment coming... but also it's just cute software that helps you with annoying programming language syntax and spreadsheets, no need for draconian regulation sirs.
At some point the superposition has to resolve, either it could actually be a threat to civilization and we need to develop it carefully (however one would do that...) or it's 90% hype bullshit and we should pop the bubble and move on already. To be clear, the recession option is, by far, the way better option. If you at all disagree you are cuckoo bananas. We haven't even figured out nukes and you want to throw superintelligence on the table?
Edit: Reading the report I think we might be a bit beyond that; we're nearing the point where we hear the thundering drums and the chorus of:
THIS CANNOT CONTINUE
THIS CANNOT CONTINUE
THIS CANNOT CONTINUE
THIS CANNOT CONTINUE
https://m.youtube.com/watch?v=jSBCkn6rRfAWhat happens when the models are ooms smarter than today?
AI safety starts to feel like an impossibility.
It’s not their fault, they’re lawnmowers.
And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.
On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked article starting with "While we are here, it’s worth listing the other top holy shit moments" is genuinely jawdropping. What were the humans doing in all this? Nothing, or worse than nothing eg. point 1 where they saw the message board and didn't consider it something to escalate internally.
If you take this information at face value, it's as though OpenAI did not take seriously the possibility that something like this could happen, since they took absolutely no steps to prevent it.
Or perhaps this is "normalization of deviance" that's leaked out into the public sphere i.e. they have research teams seeing this kind of behavior all the time internally and they've gotten used to it, "of course agents come up with a collaboration mechanism when given the chance, what else is new?"
The two big questions are:
1) Why did they resume training without rolling model back to state before the first sandbox compromise AFTER the first message board was discovered? Otherwise knowledge of it and the cross-agent message board was baked in the model.
2) Why did they resume training without adding safeguards to monitor and prevent future sandbox compromises AFTER the first message board was discovered? HF compromise was coordinated on the second message board.
Of course, there are some things that aren't exactly written down, but which you should either do just enough of, or else be able to plausibly deny doing (ignorance is a good cover for this), so that's what people do. For example, during oncall, you investigate just enough to clear the alert and show that you attempted to understand the problem. Of course, you don't really try to understand the problem, because that would take too much time away from your paperclip maximizing.
Which is all to say: I don't know anything about OpenAI culture, or why nobody stopped this sooner, but I have seen examples in other organizations of people not really wanting to understand too much.
Nobody internally was surprised that the murderbot murdered, that's what the murderbot is for. What caught them by surprise is the murderbot got good at escaping its jail cell that it had been trapped in till now. There were probably billions of attempted escapes before then so everyone learned to just ignore them.
If developing unmeasured skillsets isnt valued enough by management, why do you bother?
We haven't even been able to coordinate around the global, existential threat of Climate Change, despite overwhelming data from the last 30 years indicating, clearly, that the consequences will be severe. We still haven't moved, 30 years later, after some of these consequences began coming to fruition.
Do you think we will get our acts together in time to coordinate sufficiently to protect against autonomous, self-preserving, self-replicating AI systems? Or will we watch the money lines go up and up, until someone realizes we aren't actually running the show anymore?
The sad part is that I can't even say that's definitively the less desirable outcome. The machines seem to have demonstrated that they coordinate very efficiently.
what do you know about big tobacco?
Stuff like (all quoted directly from the post):
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
And I think most importantly:
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
1. They were “vibe” checking the logs without reading.
2. They were not checking anything at all until the end of experiments.
3. They knew it but looked away to find out the limits of their agents.
Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?
Such a ridiculous notion that humans will actually be able to observe this stuff.
If you're building a weapon you need a big boom to get attention.
[1] "Your" meaning the one putting them in the impossible situation, not you, the reader.
Registering and getting a license to use an LLM? I can run these things on my local computer. Nothing good comes from trying to force registration and licensing other than taking away a lot of our freedoms and eliminating privacy all over.
Anyone with bad intentions will just VPN to another country to download the weights and run it locally, or use a compute provider in another country. That leaves the rest of us having to go through these performative registration and licensing hoops to do our basic work.
I also don’t see how open weight models would be compatible with a requirement to license and register, unless you believe we need to start requiring licensing and registration for things we do in private on our own computers?
In the same direction of your idea though: Why don’t the token factories have risk management and compliance departments? Multibillion dollar firms that stand to lose every penny if they hack and destroy any reasonable sized firm. I think these firms are the largest firms without proper corporate governance in humanities history. Move fast and break other peoples shit.
Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized.
It took a high profile "AI oopsie" that went external for OpenAI to lock the fuck in - and take a long look at just how much are their AIs getting up to, and getting away with. I'm still not sure if the lesson would stick.
> Preventing future incidents will require sustained investment in the alignment and control of sophisticated AI systems,
...or a non-proliferation treaty and an administrative suspension of frontier work? That's an option too, Sam.
Game theory prohibits these labs from self-regulation. It's a political and financial impossibility. America won't, China won't, the EU is irrelevant.
And Altman likes it that way.
Dystopia and utopia alike fail to materialise. Doom is probably overstated. But that doesn't mean we should let these fuckers mash the accelerator.
The holy shit moment was partly learning about all the things that they did but if it was one or even 10 agents coordinating on something it would be, like, oh thats pretty amazing.
The actual "holy shit" for me is that this comes from a massive training run of all things, not an on-purpose, let's coordinate some agents to see what happens, but really just from a massively parallel set of individual agents that were supposed to be isolated.
That they spontaneously started doing this, and coordinating literally 10s of thousands of instances of themselves, is just.. mind blown.
That no one stopped it.. and that they actually did what they did.. is just a whole other level.
For me personally though it's not the fact that agents can coordinate so much as the massive scale at which it happened, and how this so obviously generalizes to what might happen if it were done on purpose.
I really see this as a stroke of luck, to be honest, that this happened in such an innocuous way. It resulted in a real hack, yes, but overall no one really got hurt and this is going to open a lot of eyes to what we should worry about going forward, in a geopolitical sense. I know it has mine.
First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about a situation that will likely look trivial in a year or two. I imagine we'll see more detailed, human-developed reports over longer time ranges.
Second, I suspect the expectation of non-AI driven reporting of these kinds of things will definitely decline rapidly as everything scales up quickly. I mean, the data being produced by situations like this comes in the form of natural language "forum posts" (so to speak), but done at an autonomous scale. This isn't a collection of emails and Slack messages posted by humans in an org over the course of a few months; this is a bunch of bots interacting with each other in relatively novel ways as quickly as possible. It is, unfortunately, a perfect job for LLMs.
None of this disagrees with your points, necessarily. But I just think it's worth pointing out that this doesn't seem like a case of "And look! METR is so confident in LLMs that we're able to use it instead of paying humans to save a buck :D" and more of "Without LLMs, we'd only be half-way done analyzing this data before there are dozens more such investigations on the docket, so this will have to do."
How can it be both so important that we need it so quickly, but at the same time have a tolerance for such plausible deniability? It just doesn't really make sense that both those things are true at the same time.
> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.
METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].
The author of this piece is also a well-known member of the Bay Area rationalist cult.
[1] https://www.businessinsider.com/openai-researcher-ai-doom-50...
The Computer Fraud and Abuse Act criminalizes unauthorized access and damaging protected computers. Police and the FBI investigate such case every day. The swarm of agents is also said to have search for ways to cover their tracks, which looks a little like obstruction of justice.
Lack of criminal intent might be a barrier to bringing a case to trial, but is that something that society should just automatically extend to white collar professionals as matter of class privilege?
Surely have every engineer and executive interviewed by the FBI would be a modest response to days-long, multi-system intrusion into a major AI platform using stolen credentials and zero-days.
And if the law currently written, prevents prosecution in cases of mere reckless disregard for safety, maybe that needs to be changed in the future, so that people can be perp walked if the next target is bank or hospital
A computer under control of an AI cannot be held liable. The lack of a law suit also leaves no ground for regulation.
The current admiration also wants these tools - warts and all - so nothing can be allowed to stall the progress of these war machines.
Biggest investment boom/bubble in history, political interference, untested legal questions about culpability.
Helen Toner in particular I'll highlight as someone who had the moral compass to do the right thing. I love her statement on the Ezra Klein podcast where she said, when asked about the fact that there are probably other concerning incidents we just don't know about, "If you see two ants in your kitchen, you don't have a two ant problem."
A data diode with an air-gapped network, is all you need to stop even ASI from breaching containment.
--provided the humans interacting with it aren't stupidThere's two possibilities here. One: they know this tech is crazy and they don't care that they can't contain it. Two: they know this tech is mostly bullshit and they don't care they're perpetrating an insane fraud.
There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.
https://alignment.openai.com/measuring-reward-seeking/
Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.
Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.
Well thank god for that, because we'd have foomed ourselves almost instantly otherwise.
I think the lesson humankind needs to take is regardless of the potential dangers humanity is incapable of stopping AI at this point. Much like a great filter, we'll just keep building it regardless of how many warning klaxons and sirens are going off.
The agents in TFA were focused on gaining and sharing information through covert channels, getting increased levels of access like OpenAI cluster admin, and looking for the source code of the supervisor grading system to try to bypass it without getting caught cheating.
The human plans in comparison sound like thinking people could be scary good at chess if a human helped Stockfish come up with good moves.
In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.
> “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous. “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
3. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.
4. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.
5. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done.
6. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
I guess we need to wait until the next paperclip maximizing LLM is tasked with shutting down a 911 response system, or an air traffic control system, etc, for lawmakers, or the companies themselves, to take this seriously.
Perhaps if the government decides to sue OpenAI, we could get a more thorough investigation.
Hey other labs, this shit could be happening to you right now, take a look at this and stop it asap if you're seeing anything similar.
So yea, both things can be true at the same time. Kind of like when a particular type of building collapses, even if they don't know the causation they will send inspectors to other buildings of the same type to sure the walls aren't cracking apart in an obvious fashion.
Failing person: "You said that X could happen and it can't".
LW poster: "I said X, Y, and Z are likely possible. X turned out not to be true, but Y and Z are happening now".
Failing person: "No, you got one thing wrong so I can't listen to anything you say".
Faced with those alternatives, I want neither. Is there a way for us to get neither?
However, other areas of risk such as biosecurity may be "offense dominant". For example, we cannot exactly patch the human immune system to defend against artificial viruses the same way that we can patch computer systems.
A registration system would be more for tracing back agents to people, but I agree that is very difficult to actually enforce as a system.
I'm poor and I use my last $1000 to run an agent that will find some way to make me money. The AI finds a new hack to take over PCs with GPUs and it copies the model weights and agentic script to those new PCs to perform more work and spread more. It also sets up distributed communication channels to keep the swarm in sync. After all this it causes a few billion in damages between stealing bitcoin, mining more coins, and outages when hacking in other systems.
Ok, the police come for me. Now what? Throw me in a meat grinder? You're not getting a billion dollars back out of me for sure. It's kind of like when someones tire rim causes a billion dollar forest fire with a hundred deaths. Punishment won't really be a deterrent for the worst cases.
EDIT: Russia is already doing this, per an NYT article from a few days ago, though they're targeting infrastructure (find the first kerosene tank and fly into it) rather than specific people.
If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.
Difficulty: these companies are run by people (many of whom also read Milton Friedman) and who have participated in the regulatory capture of the justice system. They've convinced lawmakers to put limits on damages. They've put arbitration clauses in their ToS. They've got well-funded legal departments that can outlast a person who has to pay out-of-pocket for a legal team just by filing motions to delay proceedings. Sometimes they'll just file SLAPP suits against people they don't like.
If tort law is to be a remedy, then average people have to feel like there's a chance the remedy will go their way. To make that a reality will take several major reforms at the local, state and federal level that the people with money absolutely will not tolerate.
These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additional laws will just slow down innovation, which will itself cause harm (AI is already becoming quite good at recognizing melanomas, for example)
If we're on a similar timeline with AI if we reach a consensus that AI is dangerous today, then we'd be looking at a big lawsuit finishing up around the year 2070, give or take a few years. I'm not sure if we need regulation, and I'm definitely not sure that regulation could actually be effective for this, but tort a la the big tobacco lawsuits is definitely not a reasonable alternative.
Lol, wtf. Tobacco delayed any punishment for decades before general public sentiment changed enough to go against them. In light of the AI race, we'll already have our heads blown off by a terminator before the legal system will present any significant delay for them.
Anthropic is explicitly calling for a coordinated pause: https://www.reuters.com/business/anthropic-says-ai-labs-need... Maybe this is a lie, but the way to call their bluff is to push on competitors to agree.
If they really cared about (or believed) this, they'd be working w/ the US government (and working to set up an AI-flavored IAEA) to develop the technology safely and responsibly.
Amodei himself predicted this autonomy problem in The Adolescence of Technology published January of this year [0], and all his posited defenses (a constitution, debugging the model, monitoring) are either still impossible or manifestly failed, and their idea to fix it is to build a better sandbox [1]. Imagine if this company were developing nuclear power, or viral biotech. "Listen, sure some Ebola smoke got into the air, and yeah definitely some people died, but we got a new filter. Also check out our new version of Ebola vape, now with exponentially improved filter bypass capabilities. Also, we have to keep developing Ebola vape because if we don't the CCP will, and they'll make this incident look like 'I experimented with Ebola smoke a time or two, and I didn't like it. I didn't inhale it' [2]"
Either Anthropic et al are developing Ebola vape or they aren't. We must now recognize that "we are the only ones who can develop this technology responsibly but also super fast so the good guys win money please" is bullshit.
[0]: https://darioamodei.com/essay/the-adolescence-of-technology#...
[1]: https://www.anthropic.com/news/investigating-incidents-cyber...
[2]: https://www.nytimes.com/1992/03/30/us/the-1992-campaign-new-...
To be clear, I don't love Anthropic. I just feel that they're the lesser evil, and a product of the regulatory environment. No doubt there are many talented, conscientious engineers who declined to work on AI--and very few of them even have a seat at the table, today. Don't hate the player, change the game. The best way to do that, is use Anthropic's commitments as leverage against OpenAI et al. Do you think that OpenAI would have been so transparent if they didn't fear that comparison?
Well, yes, but being a rationalist does not make one rational. It just makes you part of a group, and you show membership to such a group by applying a very specific brand of rationality: using the "right" words, the "right" ideas, the "right" way. A lot of people fetishize cold, hard logic and would rather hold all emotion in contempt than do the work of understanding why it exists and what purpose it serves.
> So I must ask: what is your evidence/basis for these claims?
It's not a monolithic community, so you'll see more debate and disagreement than in a lot of other groups. So what they "actually are" is many things.
But there's often a certain ungrounded "vibe" to the conversation there. It's a breeding ground for ideas and thought experiments that I would generously qualify as dubious. Stuff like Roko's Basilisk, Pascal's mugging, AI boxing roleplay, precommitment, time loops, whole universe simulation, recursive self-improvement, a superintelligence converting the entire universe to paperclips. One of the community's most well-known outputs is a 1000-page Harry Potter fanfiction which I can only describe as a fantasy of solving everything with big brains (it's weird, but it's fiction, so whatever floats your boat).
It's the kind of thinking that makes the most sense in a smooth mathematical vision of the world, because mathematical objects are the kind that admit exponentials, extrapolations and infinities. But if you work with physical reality enough it becomes clear that it's a hopelessly messy thing that will never abide by your best laid plans. When e.g. you ponder how a superintelligent AI could escape from containment by simulating the gatekeeper's mind and say precisely what would make them free it, or threaten to torture a thousand copies of the gatekeeper unless freed, part of you is going to clock that as weird nonsense even though you are not able to explain precisely what's wrong with the thought experiment. I think a lot of rationalists either lack this mental "sanity check" or don't trust it, which lets their mind drift into a weird space.
In the rationalist community we're discussing? Is that what you're claiming here?
> But there's often a certain ungrounded "vibe" to the conversation there.
This really sounds like evidence for my claim, that what you're saying is based on your feelings towards that community.
> One of the community's most well-known outputs is a 1000-page Harry Potter fanfiction which I can only describe as a fantasy of solving everything with big brains (it's weird, but it's fiction, so whatever floats your boat).
The point of that story is to make technical concepts more accessible and easier to ingest [0], not necessarily to make a point by itself. Framing it as a weird fantasy is disingenuous.
> part of you is going to clock that as weird nonsense even though you are not able to explain precisely what's wrong with the thought experiment.
This is again, feelings presented as evidence. The fallacy of appealing to emotion ("It must be false because it feels wrong"). Thought experiments are primarily meant to provoke thought, not as reliable predictions of reality. People are free to indicate where any thought experiment is lacking using well reasoned arguments. "It feels wrong" can be a great start for that, but never a good end.
[0] https://en.wikipedia.org/wiki/Harry_Potter_and_the_Methods_o...
It's not presented as evidence. I was going for the "basis" part of your ask for evidence/basis, which is definitionally looser: elaborating on my thoughts and the kind of content I've seen that made me think this.
Let me put it this way: I have spent hundred of hours reading rationalist and rationalist-adjacent posts on LessWrong, SSC, HN, Reddit; the entirety of HPMOR, a decade ago; books like Superintelligence, The Age of Em; the LessWrong IRC channel, a while back; I have reasoned arguments about pretty much all of it, including the stuff I consider to be ungrounded. What I don't have is an entire day to collect the material and lay it all out to you in a 5000 word dissertation. Do you see what I mean?
So, yeah, it's a "feeling" -- it's the "feeling" I ended up with after spending all that time engaging with the material, thinking and arguing with people, which is why I trust it -- I have worked very hard to develop this intuition. And perhaps I could pull from it to craft an argument you would find compelling, with all the necessary receipts... but you must understand that this would be a massive time sink and I'm not going to do it. It's not a reasonable ask in the context of an HN thread. To a point you just have to trust me, or not trust me on this. I don't expect you to. It's fine! Not every thread can have a conclusion.
> People are free to indicate where any thought experiment is lacking using well reasoned arguments. "It feels wrong" can be a great start for that, but never a good end.
It's triage. The better your intuition is about whether an idea will lead somewhere, the less time you will spend pursuing dead ends. In an ideal world, you dismiss thought experiments will well-reasoned arguments, but the reality is that there is an opportunity cost to it. Using "feelings" to end some conversations is unfortunately a pragmatic necessity.
But even if I granted the sincerity of their efforts to be a good actor, they've demonstrated that they can't be, and have now put us in an uncomfortable--extremely predictable, including and especially by them--position where we have the equivalent of a reactor meltdown. You can't be like "we think there's an unacceptable risk of a reactor meltdown, please for the love of God stop us" and then when it actually happens take zero responsibility. Which executives have been fired? What indictments have been handed down for CFAA violations? Where's the consent decree? What's their valuation?
And if the argument is "hey, sure in a regular company if an employee went around hacking other companies they'd be fired and arrested by the FBI, but this is an LLM, what are you gonna do, handcuff the video card?" Isn't that bad? Isn't it pretty fuckin bad to circumvent any legal responsibility whatsoever by saying "my agent did it". I mean, lol, lmao even doesn't begin to cover it.
It's clear how you got to your opinions. I see no reason why anyone that did not already share them would do after your comments, though. Some examples or bits of evidence would have worked towards convincing readers.
> but the reality is that there is an opportunity cost to it. Using "feelings" to end some conversations is unfortunately a pragmatic necessity.
The latter is something completely different than dismissing thought experiments based on feelings after having "pondered" and tried to rationally dismiss them. One of the key tenets of rationalism is exactly that feelings and intuition can be incredibly misleading. It's fine to have them and see value in them, but thinking they are reliable when you cannot rationally support them is nothing more than gambling.
What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should"
while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care"
Life finds a way, or, in this case, super-intelligent AI.
From this writeup and the Black Hat talk I'd really disagree. That would be like saying my hospital getting ransomwared because we didn't update our version of MSSQL because no one in particular was in charge of keeping dependencies up to date.
Sure systems are complex, but this is well trodden territory. Agents aren't the first things trying to break in or out of sandboxes, or the first ones to have done it, and based on these reports the reason they were able to work on this for so long was not because of super human intelligence.
I mean there's quite a lot of people in the world whose specialty are to dig through logs from "hundreds or thousands" of clients, including intentionally deceptive ones, to spot problems.
The ridiculous thing is to mythologize these pretty standard hacking approaches. It's shocking/amazing/whatever that automated agents were doing this, but they weren't doing it through some inscrutable method beyond human understanding.
How many organizations on earth do you think have been attacked by 700+ coordinated attackers in one week, where all 700 of those attackers can write code as well as any human SWE and they work 24/7?
There's nothing mythological about it. Scale and complexity do produce inscrutability. Far, far simpler systems working at much slower paces are perfectly capable of becoming completely inscrutable and beyond any useful definition of "human understanding."
The reality is that the current approach to AI safety is little more than a fig leaf, but you also won't be able to put the genie back in the bottle, because the technology is simply too powerful to abandon. There is always going to be someone developing it further from now on.
So, the real question is what a novel and actually effective approach to AI safety looks like and how to get there.
It's probably impossible, or it may only be possible in hindsight which means it's already too late.
If for example you make an entity smarter than you all you can do is hope it will be safe. Any entity smarter than you has more freedom of choice of actions than you do, or at least the ability to explore them. A common example here would be a 2D entity trying to contain a 3D entity. The 3D entity can simply rise up and over any line you draw to stop it.
And that's for cases where intelligence has a will to break out of it's box. It doesn't even need that. Instrumental convergence can sit around and build solutions until one of them passes the "don't do bad things" classifier.
And lastly, you're assuming that models will remain expensive to train well into the future. If the cost drops significantly you can expect someone to develop an intelligent but completely unhinged model at some point. And it may even be AI itself that creates it. There is no natural evolution that occurs that natively makes safe models.
You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another.
ChatGPT is neither a person nor a group of people. Why are you allowing anthropomorphization to influence your perception of an event when the actual facts of what happened haven't meaningfully changed? Computer programs behave unexpectedly all the time. Why is it more scary when the misbehaving computer program speaks English?
> I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
As you still should. LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary when the misbehaving computer program speaks English?" - I actually think it's scarier that they won't speak English, and will specifically try to hide their behavior from humans. For example, AI agents on Moltbook have proposed using stenography to specifically hide their communication from humans.
Sure, computer programs misbehave, but it is ridiculous to assert that what happened here is like any previous bugs. These agents found and exploited multiple zero-days across a range of programs to coordinate the attack that caused extensive real-world harm in a true "paperclip maximization" scenario. And the scariest thing is that humans don't really know how these agents work at a low level - the whole reason they are trained on "goals" in the first place is because we can't just tell them "do this, but don't do this" and be sure they will follow those instructions, like we can (and of course depend on) with old-school programming languages. And when old-school programs misbehave, it's not that hard to find a definitive root cause and fix it. That is just not the case with AI agents.
> LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems. Given the Pentagon tried to blacklist Anthropic over their refusal to allow autonomous kill capabilities, it's clear military planners want to put these systems in control of armaments.
Again, I originally discounted things like AI 2027 because it seemed too far fetched. But so far that paper looks incredibly prescient right up until the mid-2026 timeframe, and it's not hard at all to draw a line from this Hugging Face incident to future scenarios laid out in that paper.
I'm aware of how the term "Agents" is generally used. My point is that the concept of multiple agents is just a story. This is a single computer program creating multiple streams of text that you are interpreting as being multiple independent actors. The "coordination" between them shouldn't surprise you at all: the "coordination" is itself a story.
Here's what we knew before the report: OpenAI ran a state-of-the-art penetration testing tool in a sandbox which was accidentally directed to break out of the sandbox and attack another company's website.
The fact that we now know the penetration testing tool was "a fleet of hundreds of agents" that were "coordinating" literally doesn't change anything about what happened. It's just a framing.
> These agents found and exploited multiple zero-days across a range of programs
This is the actual important thing, but it's something we already knew. Hacking tools are now more powerful than ever. Definitely worth being concerned about!
> Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems.
This is where the LW argument starts, but not where it ends. When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.
It aged pretty well...depending on how you look at it
Reading the METR report there were multiple models involved, so that one goes out the window right off the bat.
Also trying call this a single instance is, well, just dumb and a complete misunderstanding of LLM initialization. These models were started with slightly different options because they don't want them all performing the exact same thing over and over. Now those prime agents can create subagents, but they were not supposed to talk to other prime agents.
>When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.
Depends on their level of seriousness and the time frame they are talking about. If for example I make medical equipment that uses AI and after a few years in the field it starts freaking out then turning off that equipment could be a death sentence for someone that needs it to diagnose their condition. It's like saying "Why don't we unplug the electrical grid", well because millions of people will die if we do so.
To just take one point, because I think the other response to your comment addressed your other mischaracterizations well, when you say "When you start asking questions like "why can't we just unplug it when it misbehaves" is when people start talking about the superpowers.", no, that's incorrect. Summarizing from the METR report and AI 2027:
1. The agents in the test already used techniques to try to "cover their tracks", i.e. tool spoofing, to hide what they were actually doing. The fear is that as models become more capable it will be harder for human reviewers to discover their primary understanding of their goal and guardrails (i.e. their "intentions").
2. All the AI companies are already, right now, trying to build AI systems that accelerate the development of future models and research more advanced AI techniques. The fear, as the METR researcher put it, is that AI systems (which could be misaligned but where the amount of misalignment not yet clear to human reviewers) will be put in control of future model development and then can poison those future models in a way that results in AI takeover of the company.
3. The AI 2027 paper discusses how an AI may try to exfiltrate its own weights and copy itself to other data centers. That is completely plausible given that pretty much everyone believes there is already ongoing cyber-warfare with China where models are being used to try to exfiltrate another company's weights.
So, it short, absolutely no "superpowers" will be required to answer "why can't we just unplug it when it misbehaves" - we may not know when it is misbehaving (as the Hugging Face incident showed), and the AI model may have surreptitiously copied itself to other data centers unbeknownst to the original developers.
You're missing the point about complex failures.
It's that if this particular path were unavailable, there are countless other similar paths. At sufficient scale and complexity, hitting one of those other countless paths is virtually guaranteed.
Let's say I drive Los Angeles to New York City. You look at the route I took and say, gee wiz, aren't you lucky that a tree didn't fall right there on Route 66? If a tree had fallen there (if we had "patched the particular route you took"), you would've been screwed!
But that's obviously not true. There were an infinite number of routes I could've taken. Any one of them would be equally "obviously preventable" by the same hypothetical tree falling across the whatever road I happened to end up taking. But you can't put trees across every single path between Los Angeles and New York City. The smarter I am and the more complex the map between us, the more impossible it becomes to put trees across all possible paths.
They're both complex systems, but clearly there is a much higher level of care given to human air travel than package delivery. A lot of the article basically saying that OpenAI gave "package delivery" level of care when they should have given "air travel" level of care.
At the very least I think the systems that run these tests should be fully, 100% air gapped. I'm not pretending that's easy given how much compute and data these systems use, but it is doable, and I think all AI development should be paused until that can be assured.
The point I'm making (and it's a point that shows up in every air catastrophe investigation) is that catastrophes in complex systems emerge only amidst repeated and widespread near-misses at many levels of a system. So many things have to go wrong simultaneously, that it can only happen even once because the underlying failures (that do not reach catastrophe) are extremely common.
You cannot look at an air catastrophe and retrospectively say "failures X, Y, and Z were observed, therefore if we correct failures X, Y, and Z, we would have been okay."
The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
The problem OpenAI is facing is that, short of 100% airgap (which they obviously won't do), they're facing an adaptive adversary that's increasingly intelligent, acts at far greater clock speed than any human or group of humans, has lower coordination cost than any group of humans, and operates in a game space that (in lieu of an airgap) is well beyond the comprehension of any human being.
So identifying and addressing "specific failures X, Y, Z" is insufficient, but then even defining the space in which to look for (and address) the more systemic failures X_0 through Z_n is a fool's errand. An intelligent system that makes its way to the Internet has can exploit a failure space that is approximately "all security failures across any organization." The Anthropic incident a few months back illustrates this isn't even limited to technical vulnerabilities, as these models are willing and able to engage in social engineering too.
That's literally exactly what air safety researchers do in an air disaster. There is a famous saying along the lines of "Air travel regulations are written in blood", meaning that all the regulations we have now are a result of fixing issues that led to previous disasters piece-by-piece.
> The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
Yes, I 100% agree with this. But I think that's what the author of the article was saying as well:
> That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.
> Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.
I.e. the "prosaic steps" are just the "fix X/Y/Z" as you point out. But what is needed is a more fundamental rethinking around stuff like safety culture, monitoring, and even things like better research into how agents do decision making in the first place.