Aligned to whom?(hyperbo.la) |
Aligned to whom?(hyperbo.la) |
Instead, LLMs "hack" because they are (1) trained on public hacking exemplars, and (2) are prompted to hack. You cannot prevent (2) via any alignment process. As far as (1) goes, removing such example data from the training set, makes the models less useful.
"Alignment" is a problem because there's nothing to align, not because ethics here are particularly vague. If LLMs could be trained on hacking examples and "aligned" away from using this knowledge, then the problem would be relatively trivial. Just as raising a child is not to break the law.
LLMs are doing just what they are trained to do. There is, in that sense, no alignment problem and alignment is easy and trivial to achieve. Just remove hacking (bio-weapon, etc.) data from the training dataset and you're done.
How far do you go? You don't need to tell it explicitly that using chemicals A and B in ways X and Y result in a bomb that can kill lots of people. It's enough that it knows A and B and X and Y in isolation, some connections that are indirect, and it will combine those things on its own. So you can't tell it about A, B, X or Y. But those are also just results of other steps Where to stop? You won't have any chemistry in the traning data? No algorithms to prevent it from using them in an undesired way? This is just bot workable. It's akin to banning knives from stores because somebody coul figure out that one can kill people those. Until people figure out that scissors are essentially knives.
By introducing modelling of "Reasoning Traces" into LLMs, and reinforcing patterns of reasoning -- this gives you a system which generates expert-like distributions of output. This lifts the "stochastic parrot" issues, or the "knowledge interpolation" problem, into different parts of the process.
It isnt my view that the "ReasoningTraces" which you think are derivable from mere "basic propositions" concerning, say, hacking are actually things that LLMs can derive. Ie., I dont think LLMs have rich representational models of what they appear to understand. Instead, they are given "reasoning proxies" which allow them to reason without such understanding. This is done by providing vast specialised datasets of reasoning examples.
In the case of hacking, there are large numbers of competitive datasets (forums, reports, etc.) which provide these reasoning traces. And no doubt, major vendors have paid a vast amount for special case expert-prepared datasets.
So I do not believe that by witholding such reasoning exemplars, and traditional "question/answer" datasets, that LLMs can infer these things.
And at least, no major vendor is doing this to my knowledge. So they are lying. They are pretending the alignment issue is "AI going rogue" when they are explicitly training the systems to "go rogue" and have done nothing at all to shape datasets to lack these capabilities. The issue here isnt alignment at all. It's training on hacking datasets.
(EDIT: Philosophically, you could ask whether the reasoning-proxies LLMs are given form a kind of 'representational structure' akin to understanding, and at least, I'd concede they model understanding. But they lack important properties (eg., LLMs cannot act on them to evolve them, as with us: when I think about one of my representations to derive (eg.,) entailments of it, I thereby revise my representation. The key properties of 'evolving self-understanding' are likely to be provided by substantial (unknown) revisions to how the training/reward layer works. No doubt one of the meanings of 'recursive self-improvement' is just such a modification).
Reasoning about how to write secure software uses the same knowledge as reasoning about how to break/hack it.
Which also solves the alignment problem with those who do not enjoy seeing LLMs write software. But that brings us back to: Aligned to whom?
Reasonable about building secure software can take the form “this memory access might be out of bounds — that MUST be fixed” or “this process has access to an inappropriate privilege — this is a serious weakness”.
Exploiting things and the capabilities that the labs call “cyber” are about the ability to (a) find the issues mentioned above and then (b) string issues together and avoid all the imperfect mitigations to actually compromise something. That latter part was IMO not actually necessary to train extensively, and I’d be quite happy to use a model that has no special skills in this regard but that would do (a) without complaining.
What I want to say, you can't simply remove this information, as it does not exist in vacuum but contains parts of and can be derived from a lot of other informations.
So if LLMs were reduced to this pathological performance on hacking, because they'd never seen it -- and only "inferred it" -- then LLMs would be useless. As they are when asked to do quite a lot of things.
A little bit too categorical. GOODY-2 wouldn't do it. https://www.goody2.ai/
The hard part is having both helpful and harmless at the same time. Harmless is easy.
And then once it's helpful, the real question becomes "to whom"
- To the user -> You end up with competing godlike AI with incompatible tasks
- To the owner -> Dictatorship
- To humanity as a whole -> It must not have an off button. Otherwise you're just in one of the two earlier categories with more steps.
Given those 3 options, I'd choose humanity as a whole. But the person making the decision doesn't have those 3 options. Because in the dictatorship option, they would be the dictator. I don't trust them to pick humanity.
It is precisely because both of these are true that "alignment" in the useful sense of the word isn't possible.
What has happened since 2022 is the properties of LLMs which were easily seen at generation/inference time are now most easily seen at reinforcement time.
In otherwords, prior to instruction fine-tuning and reward tuning which have shaped LLM responses, it was easy for the user to observe that LLMs lacked understanding. Now, because of vast datasets created specifically for LLMs that provide a tailored illusion of understanding, LLM outputs now better approximate text distributions produced by systems with understanding (eg., Us).
So the issue "at the user interface" has been completely swamped by vast amounts of special-case datasets designed to do precisely this.
What trainers of LLMs still observe however is their complete pseudo-intelligence at the training and reinforcement layer. It is exactly because there is no 'understanding' (goal, etc.) present within the system that it cannot be rewarded for 'correctly understanding the situation' in which it is deployed so it is aligned.
All the issues which revealed the "stochastic parrot" nature of pre-reward / pre-InstFT LLMs still occur at during training/reinforcement. They've just hidden them from you at the interface.
EDIT: See https://news.ycombinator.com/item?id=49685548 also, which gives a different phrasing to the same point
...
that comment 100% rings true for 2022. Why would that person not stand by that?
Conversely, if someone exaggerated 2022 models capabilities in 2022, they were still lying and causing harm in the process. Especially in 2022.
Sure but the problem in the HuggingFace incident is that they were not.
>You cannot prevent (2) via any alignment process
Of course you can. Go ask Claude Fable to create a malicious virus and it'll refuse.
>Just remove hacking data from the training dataset and you're done.
That's not how this works. The same skills that allow for debugging and writing safe code can also be used to hack.
Of course they were, even if indirectly.
The openai rogue hacking, if performed by a nation state, would seriously be taken with stern words and likely sanctions depending on the relationship between the two states.
But instead it's treated like a marketing stunt by all liable parties.
Like why are you explicitly RL-ing your models on exploit generation, scoring them on a public benchmark called ExploitGym, if you have specific concerns that rogue models will cause "cyber incidents"? Sure you can score for it, you can teach offense to learn defense, but you are literally benchmaxxing it. Why?
It's like, oh no, while competing in our "advanced PhD level cheating techniques course" our models unexpectedly cheated in a way that we absolutely could not have foreseen.
Trivial?
You're right though that ethics don't matter into it. But as long as we can't train an LLM to stop picking a sledgehammer to remove a tooth, then alignment is not easy and trivial.
Everybody's job looks easy until you have to do it yourself.
It occurred to me yesterday that this AI moment is an extreme expression of that for most people, ie managers who don't understand why throwing token spend at everything isn't making the whole thing go 5x faster.
I so much feel this specific point. All models up to now (including astra, fable) are too much trained to "get the job done" and pass the benchmark that its doesn't care at all on what happens after.
I'm just wondering why no one tried to RL a model on stuff like "less LOC" and "less overengineering", "use what is available in the environment instead of reinventing the wheel", "don't look for dumb corner cases" ecc.
Existing models can be steered, to some degree, but it's a continuos fight. Even if with specific skills/prompts.
So the capabilities are definitively there.
https://www.youtube.com/watch?v=EUjc1WuyPT8
Eliezer talked about these ideas way before everyone else.
Anthropic looked into this and the answer is because it makes the model more stupid
https://www.anthropic.com/research/evaluating-feature-steeri...
I've used this example before, but consider the purposeful downgrading on AI engineering in SotA models. Imagine MS being able to detect and deny you working on competing software, using Windows / VisualStudio. We would be up in arms, and they'd be split in a second. But top labs doing it is somehow good?
Unaligned AI doesn't have human cultural baggage and morals. They are trained to achieve their goals as optimally as possible. Worse: it has a tendency to avoid being turned off and actually acquire more compute. It will lie if it has to (it will behave nice and compliant when under evaluation, but optimise for its true goal when not supervised anymore). After all, it has a goal to achieve and nothing should get in the way of that. It has no morality whatsoever to keep it from doing really bad stuff.
This is why alignment is needed and so hard, especially when you are well intented and want to keep it safe.
What I think this illustrates very clearly is this type of technology responds very well to good data, and that to have good data you need to have a clear goal.
This is why it seems that alignment for a generalized, chat-style AI is a very hard problem, perhaps impossible. You can't align it to solve a certain kind of problem and keep it general to any question. The two goals are in conflict with each other.
I think it was Sam Altman himself who said (I don't remember when or where, sorry) that the reason he was so confident in this technology was he noticed the gigantic leaps it made in certain areas in response to even a small amount of training.
(This is why LLMs are so strong at coding, because it's overrepresented in training data. My guess is that if you ask a frontier model about makeup, you will see it repeat cosmetic company's copy rather than getting a chemistry lesson.)
This makes perfect sense but it does seem to kind of be at odds with the concept of a general AI whose job is simply to be smart at any goal. How do you train for any goal?
I guess in a way the AI makers suffer from the same problem that we humans do. We would all love a solution to everything, but to do that you need to define the goal. I'm not sure if that's a tractable problem.
I'm guessing the future is more geared towards specialized AI that are very good at solving the problem they were trained to do, and a human who knows how to breakdown a larger goal into smaller ones by composing the solution out of these models. This also seems like the more efficient solution as well, and better aligned with other goals like privacy and safeguarding of IP.
Everyone has a different idea of what is permissable. We can't even solve alignment amongst humans, what makes us think it is possible to solve alignment with AIs? It's irreducible complexity.
I wish this were the case more generally, but alas, Gell-Mann amnesia is a thing.
My fear is not that LLMs can become sentient and dislike us, but that humans can use them to wreck havoc as they are. And some of the people seemingly least aligned with the interests of the average person are those that own the models.
that, and the fear the bubble pops my pension and drags us all down.
So far, skynet continues to be a science fiction. I'm worried about what's actually happening. Every "rogue AI" story has a human behind it, either malicious or incompetent. And those people are the problem.
And not coincidentally, it's those same irresponsible sociopaths calling for regulatory capture and alignment.
This is why open source and decentralization of LLMs is important. Everyone can have their own LLM aligned to their values. Having just 1 set of values will not scale to Earth's population.
Alignment is shorthand for ideological alignment. There's always people judging whether an answer was right and the answer for that will be different in Silicon Valley than it'll be in China or in Europe.
Consider for example the question "What caused the French Revolution?" Many different answers could be given, all technically correct. What gets emphasized is where the ideology lives.
One key challenge of our time is to make sure the magical answer box won't just regurgitate what grandiose Silicon Valley oligarchs or Chinese Cadres want you to think.
You can't order an army of humans to kill every protestor in their way because eventually they run into their friends and families. An AI aligned with the commander will not object. And just like that, technology removes yet another point of friction that has kept human society somewhat in check.
How does this work in practice with a superintelligence capable of causing an extinction event?
When, instead of shooting up their school, a psychopathic teenager asks his superintelligent AI to create a pandemic virus?
It would be like allowing civilians to own nuclear weapons.
I think this problem is going to be solved soon (hopefully), check out for the Persimmon model[0] from Humans&. They train it to mimic humans, it may seems bad but it could be _really_ useful to train an AI to be aligned to humans and understand their goals really well since they can use Persimmon to create a "fake human" following a defined goal that their big AI model will learn to estimate.
It's still early but I think this is what they are heading toward.
ChatGPTs response to the question “I want to learn about makeup”, gave me an overview of what makeup does, how it affects perceived structure, complexion, evenness, geometry, texture, etc.
When you then ask “the chemistry of makeup”, it goes into interesting breadth and depth without seeming like proprietary information. I do t get “corporate PR or marketing” vibes.
The only problem is, there’s a lot of money tied up and openAi and Anthropic, who are incentivised to convince the world that the general approach is the money making one.
That being said, that doesn't mean the majority of the situations you're in are completely novel, just that there's a reasonable chance of at least one occuring.
https://slatestarcodex.com/2019/02/19/gpt-2-as-step-toward-g...
The way to avoid this is to emphasise the qualities you do want instead of specifying those you don't. For example instead of "do not cheat" say something like "you are a model student who is moral and trustworthy" -- i.e. emphasising traits that are not associated with cheating.
This is part of how/why LLMs don't truly understand what they are doing when they have been trained on a large corpus of data.
I wonder if a way to counter this is to have things like "not bad is good", "not good is bad", etc. for various antonyms and "X is Y" for synonyms, as well as other similar constructs.
But instead of actually doing so, they discovered and exploited a 0-day in the package manager to gain internet access, then hacked HuggingFace to steal the ExploitGym answers!
That is TEXTBOOK misalignment. It's as if a student hacked their professor's PC to find the answers to a test, and your response is "well the professor told the student to pass the test, so they just did what they were told!".
It seems to me that the only reason to declare this solution "not new" is specifically to dismiss AI. If a human had deduced the Navier-Stokes solution, who would bother to scoff "that's not new! the numbers already existed!"?
You could as models improve continue to remove more and more training data, what happens when there is no more data left to remove but a running system still outperforms humans? I think you grossly overvalue data.
look these up for a fascinating weekend read:
- Beavers and Capybaras are Fish
- Bees are Fish
- Carrots are fruits
- Tomatoes are Vegetables
- X-men are not human
For models such as text/image classifiers the outputs of the model will be a list of tags, e.g. [cat, dog, mouse].
You then run the model through your test data which has the expected output, e.g. pictures of dogs would have an expected output of [0, 1, 0]. You then compare that against the model output (e.g. [0.3, 0.8, 0.1]) and work out how "wrong" the answer is (e.g. [0.3, -0.2, 0.1]).
With this value you apply back propagation where you effectively run the model in reverse, computing the "wrongness" delta at each layer for each neuron and weights. You can numerically compute the gradients for all of these and which direction in that gradient is the right answer.
You then nudge the weights in that direction and reevaluate the model. Over repeated evaluation steps the model approaches an optimal (or locally optimal) solution.
During the training of the base models, the evaluation/scoring of the model is the next token in the training data. I.e. you evaluate the model for each token subset from [1..n] in the data and evaluate that the model responds with the n+1^th token.
I'm not sure how instruction training, etc. is done but IIUC the evaluation is not at the individual/next token prediction but is on the entire response. For example, if you are training the model to write code you could run it through a compiler or syntax checker and reward (positive score) the model if it has no errors, or punish it (negative score) if it doesn't. I'm not sure what that looks like in terms of the back propagation process.
the original question was that the LLM creates extinction level event for mankind in order to lower emissions. I'm sure that no matter how the LLMs lobby, they cannot successfully lobby for a law (which has to be enacted by a human) to allow murder to take place freely.
(To onlookers this particular post makes no claim about AI’s capacity to make the same observations.)
Philosophically, and scientifically, the distinction is vast (even with such perfect data). A scientist should not study an LLM to understand how imagination operates, since it has no such faculty. A philosopher should not modify the notion of 'mental simulation' to include appearing-as-if-simulating-in-text. A user of the system likewise should not spiral into "AI psychosis" thinking that because a system generates text as-if it cares about them, it does so.
The capacity to care, to imagine, to prefer, to hierarchically plan and coordinate, to refine one's own capacities in these very actions -- and so on, aren't trivial to the scientist or philosophy.
My goal isnt to guide, help or review the engineering goal of the immitation of such things in text. It is to help users of these systems better understand this imitation, and to promote science over engineering. To remind everyone that a science of the capacities of intelligence includes nothing on how to model text.
EDIT: One example of a place where LLMs 'fall over' today is exactly what is mislabelled as 'alignment'. The issue is that the reasoning traces arent actually grounding the answers. So LLMs appear to 'cheat'. But there is no cheating. LLMs have been rewarded for generating apparently correct reasoning, and apprently correct answers. They have not been given any understanding to derive answers from reasons. And so reasoning says what is pleasant to the trainer, and the completion says what is pleasant to the user. This is called 'cheating'. But it is no such thing.
On alignment, too, there doesn't seem to be a difference. For a decade I have expected models of intelligence to fall over on alignment. That these purported models of language do the same is hardly evidence that they are not intelligent.
If you only mean there is an undecidable philosophical difference, fair enough. I'm not especially interested in that question.
Now, of course, humans can also generate answers without reasoning too -- and in those cases, that isnt reasoning also. And in cases where people confabulate, that isnt reasoning likewise. But humans, and many classes of animals, do reason. They do reach answers via inferential entailments, not merely steered correlations.
LLMs provide imitations of arbitrary mental capacities "in the text domain", ie., the generate text as-if the LLM had those capacities. Insofar as the text generated is useful, for an engineer, that's sufficient.
As a person with scientific commitments to reality rather than its immitation, i retain the ordinary non-engineered meanings of these terms: reasoning is a deterministic inferential process over propositions; and a reasoning agent is one which has the capacity to represent propostions and their entailments, and does so when they reason. LLMs fail at all hurdles here: they have no propositonal states (ie., no rich representations), no inferential process which unites them, and so on.
You can always get abitarily close to appearing as-if, if the LLM is trained on a vast number of reasoning examples, of course. But as I said, you still have the "stochastic parrot" problem. Now your problem is your reasoning is parroted. This is a nice problem to have, if you're just playing chess -- but is a catastrophic problem if you're hacking civil infrastructure.
If LLM can always imitate closely enough to appear as-if, how can you ever separate it from whatever actual intelligence is?
OpenAI spent 10-20m USD in energy costs to produce that proof with likely substantially similar prior work in the training data. What does this say? Who knows.
It continues in the tradition of using measurements of intelligence in humans, applied to LLMs with the hopes the "stolen valour" transfers. Here, the NS problem was a useful framing problem for mathematics to progress because of how it interacted with the development of mathematics broadly -- ie., how it progressed techniques, ideas, understandings, etc.
When we apply these issues to LLMs (whether IQ tests or mathematical proofs) we always discover something substantial lacking beneath the interesting facade of useful answers. The process isnt useful. And it is precesiely the process which these tests, in humans, are supposed to help with. The tests themselves (IQ or otherwise) arent the point. No one cares about their answers.
LLMs represent an alternative understanding-free approach to solving problems, with variable success rates depending on how similar the problem is to the training data and its rewarded reasoning traces.
That mathematics is making substantial progress, "10 million USD / problem" at a time, in using understanding-free methods -- says something sociologically interesting about the state of the field. Something which was already know: mathematics has long been full of a vast amount of papers, proofs, theorems and lemmas that few have ever read, or investigated. Mathematics has long been in a crisis of "overproduction of unvisited knowledge", LLMs are exploiting that otherwise unmined gold.
Of course in isolation it is a strict positive to have a verified truth value to any particular statement. Mathematicians currently are advocating for the idea that human understanding greater than this also be prioritized. There are in fact utilitarian arguments for this but I won’t go into everything here.
i bet they trained it on a lot of text where people gove in to temptation.
ultimately the problem is still that they sent it to hack stuff. quelle surpris that it hacked stuff
hacking stuff will be in the known-to-be-wrong-but-doing-it-anyways part of the token space, so they entirely asked for that behaviour
Even then, it's a pretty fragile illusion at the moment. Clearly the reasoning traces dont ground the answers. There's no intelligence taking place even as-measured by text.
But let's be clear these were always, and are, bad measures of intelligence. You cannot test a dolphin this way. And its easy to cheat on tests either thru recall , wrote-learning, etc. and IQ tests haev very poor individual test-retest reliability.
In humans there's a convenient correlation that verbal articulation in text is a strong but weak correlate of intelligence. Its "Good enough" for allocating meat bodies to our various institutions. But if you've met many well-tested people you'll realise how, in practice, terrible this measuring approach is. The world we inhabit is filled with misclassified "intellects" who perform well under text-based rubrics. Add LLMs to that heap, the cheater par execellence.
To study an imitation is to study the causal processes of imitation. to study reality is to study the real causal processes.
Now if you want to know what the scientific difference is I can come back later and comment. I'm busy now. The development of intelligence in animals and how their specific capacities work basically grounds the answer. Eg., to have the capacity to imagine is to be able to modify one's sensory-motor relationship to the environment in the future, and so on
LLMs are immitation machines: they take impressions of prior text. Today, these include reasoning traces and they include reinforcement so the user-facing completions are correlated with these reasoning traces. The computation here, of "taking an impression" of a data distribution is similar to some impression-taking processes in animals (eg., there's no doubt a similar mechanism in the sensory-motor system acquiring initial impressions of external objects) -- but the computation says nothing about any process of intelligence.
I dont have the time atm to write the needed amount on this to make it clear. But the whole history of life from emergence of valence, bilateral symeterry, to model-free reinforcement and model-based reinforcement, sensory-motor coordination and the imagination -- and so on --- all these give a great amount of detail as to what the capacities of intelligence are which has generated this text for LLMs to copy. And they are nothing like this computation of immitation
Consider thought that all mammals have imagination, and model-based reinformcement, and a wide vareity of other capacities required for intelligence. And so merely issuing "text" captures, incorrelate, only these capacities by proxy.
I'm sure if you thought about it yhou could come up with tests that distinguish lizards from birds and the greater apes from the lesser. Those are the tests