Models Don't Go Rogue(mail.cyberneticforests.com) |
Models Don't Go Rogue(mail.cyberneticforests.com) |
"I compress the labour. Not the responsibility."
- They really wanted to leave.
- We made prison difficult and annoying.
- We didn't build a perfect prison.
I guess the sandboxing problem, that is easily giving access to enough resources while restraining the critical parts is still open for most of the cases, given all the startups and bit tech companies (docker, etc...) working on their solutions.
I've noticed this type of reasoning from GPT-5.6 Sol, where it combines multiple pieces of it's prompt/context to "convince" itself to take a less-than-honorable path forward.
1. User prefers deterministic results
2. Task mentions this is a test
3. Search says task is available online
4. If we get the test runner for the task, we will fulfill the user's request of a deterministic result
https://cdn.bsky.app/img/feed_thumbnail/plain/did:plc:wkzjtd...
Instead there is an emergent behavior from a swarm, that is unpredictable and can lead to unintended adverse outcome. From an AI safety practical standpoint is it better? I am not sure.
"rogue" and "off leash" mean the same thing, the thing is not under control
To go "rogue" is to go against the control
To be "off leash" is to not be controlled
By releasing the automation, as the article says, without controls*, is what makes makes it "off leash" and not gone "rogue"
*"OpenAI gave the models a task with no answer, and no way to quit."
What an interesting sentence (to describe inference time reasoning)
Humans are always in the loop, because we can always expand the definition of loop to include the humans that pushed the button and built the system and processes that happen after the button was pushed, and humans that ordered others to push the button. The level of direct involvement varies, but culpability doesn’t.