This one seems possibly PR because it states what the others would have stated were they (good) PR: “the model had the power to hack in, but it was wise enough to not be evil”, and stopped, and left the network untouched, and didn’t cheat
Corporate America will love this one, while being petrified of the others.
Maybe it’s true though. Doesn’t really matter at this point.
Yes, the models are smart when they find a way out, but their instructions are, in a way, deliberately open such as to be a case for misalignment in these cases anyway. It’s a capture-the-flag assignment in the Gemini case, and in the OpenAI cases they were broad instructions to best the reward function. As part of alignment studies, this is literally what you’re trying to observe and then work with. If Irregular’s sandbox had been a little more boxy, there wouldn’t be these issues.
I’m not saying we have zero problems here on the AI side, but I’d certainly be reviewing my contract with Irregular at this time if I were playing in the space. It would be just as interesting to learn more about their sandboxing techniques as it would the models in these particular scenarios.
The alternative explanation is that alignment is really so bad that they can't prevent it.
Either way, all of the major AI players should be embarrassed and held accountable. If humans did this kind of thing and got caught, they'd go to jail.
There are many AI saftey researchers that have been around from long before LLMs that talked about how alignment may be completely impossible in a general intelligence agent. Look up their work from before LLMs.
We've watched milestone after milestone of their warnings get hit. It would be like finding a book that describes everything in your life. And as you turn page after page you're in a chair reading the book you are holding in real life. But you look and there are more words. They are future words. And it's getting quite worrying because there are only 2 more pages in the book.
As somebody who implements LLM tooling at work, in my experience it makes people a lot more skittish and demand a lot more in terms of safeguards.
Why does anyone use this company?
Real life is like reading a fiction story where a group of people are trying to prevent an unfolding disaster. The crazy thing is they already have the instructions on how to prevent the disaster. And they go on to ignore the instructions and with every step make the problem worse.
Most people would put it down as being too cliche.
(And we’re fixing many of these things, but worth noting this happened at a third party vendor, not in our lab)
They made mistakes, obviously, but people are so conspiratorial these days that they just assume unlikely things off the jump.