even vending operators understand the idea that a machine left in service while in deficient state, is a lose lose scenario. if you let it go empty of change, or product, or dispense substitutions whithout prior warnings, the user looses a buck or two, the operator loses substantially more, thus the service maintenance model of extract and replace, rather than continue to vend in a default, saves money and preserves chattels.
To repeat a point raised on Reddit: wouldn't this imply that Anthropic is involved in slavery?
My opinion: https://news.ycombinator.com/item?id=50012423
Often abuse comes from a bad place anyway. As the programming meme goes:
> do what I want, not what I said
which can create horrific alignment issues with conflicting demands.
How would you know, given you never get any other kind of output?
Check out spec driven development. If you constrain the available search space by being very specific it works very well. If you leave everything open and leave it to make its own decisions YMMV.
I'm a big fan of #Need as a top level heading to let the model know _exactly_ the scope of the work and its purpose, as this can mitigate alignment issues.
It is easier to stomach when treated for what it is.
Just more false advertising.
I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting". It's the whole game. Either you can trust it or not. Every other way quickly devolves into a whack a mole guardrails game with very early diminishing returns, for very good reasons.
Now if they kept writing in a very misleading way or in broken English though, or omitted crucial unknowable context...
People who are not particularly online tend to be terrible at writing messages of any kind in my experience, including agent prompts. They fail to properly distinguish what's obvious and what isn't, and how things come across tonally, simply because they're not used to the medium. So this I can imagine.
Sure:
https://x.com/lifeofjer/status/2048103471019434248
His prompting is likely confusing the agent. "NEVER FUCKING GUESS!" is a terrible prompt, combine it with "NEVER run destructive/irreversible git commands (like push --force, hard reset, etc) unless the user explicitly requests them" and you can easily get issues. That "unless" is awful and the idea that a model can tell if its guessing or not is very non-trivial.
Just generally the language he uses and his attitude demonstrate that he's not the sort of person that should be doing this thing. He's not skilled enough to understand the risks.
> I don't see why giving an agent the access necessary to perform actions would be "unhealthy prompting".
Sorry, two different issues. The major factor in this happening is IMHO:
* Don't give agents the keys to production. Instead have them write tooling that you can test and you press the button in the tooling. No AI, deterministic process.
* Abusive prompting - this can create alignment issues because if you're angry then it might conflict with previous input, this can make a model unsure about what to do and act in unexpected ways. Also it creates the "don't think about a duck" problem.