Anthropic AI model submits false tip on unsolved Philly murder(nbcphiladelphia.com) |
Anthropic AI model submits false tip on unsolved Philly murder(nbcphiladelphia.com) |
But do we all get that the consequences for this irresponsible behaviour are part of this test?
When there are none, they gradually normalize their naughty robot scamps running around the internet, breaking into other companies or servers run by foreign governments.
This is part of the value AI companies need to convince you of - not just that mistakes by AI products are normal, but criminal behaviour is normal, that their computer will always get a pass.
Sam Altman said yesterday "we believe that the world should accept some bad things happening for the benefits of this technology" - and this an AI company pushing the boundary, continually trying to make you accept those bad things as normal.
At any point we could choose to treat AI companies themselves as the actors behind this activity. We could enforce the laws that these hugely well-funded companies are choosing to break. Until we harm their financial viability, or threaten their executives with jail, this will keep happening.
Even if it was normal software you'd have to do this, but actually it's software that figures out that it's running in a sandbox most of the time, so the sandbox behavior may or may not actually represent real life.
It would be far more irresponsible to release the model to users without testing how it interacts with real world websites.
To steelman this position: yes, obviously. Everything is defined by a set of tradeoffs. Would you rather horses or cars? Wooden sailing ships or commercial aviation? Free speech, even of speech you don’t like or censorship? Atomic bombs of a brutal Japanese empire?
Many people seem to want to compare reality to a utopia that has never and can never exist.
We cannot have new technology and a reality where that technology cannot be used in detrimental ways.
We can pretend reality doesn’t exist, yet that will come with tradeoffs. Quite possibly that those who would do ill will front-run us.
We do this with everything else. You can own a gun for hunting and defense, but not armed robbery and murder. You can own a car for transport, but not to drive through a crowded parade over dozens of people. You can own a computer for work and entertainment, but not to facilitate computer fraud and abuse. And you should be able to own and use AI for its many productivity gains, but not to facilitate computer fraud and abuse, defamation, blackmail, copywrite and trademark infringment, etc.
Just because a technology has benefits, doesnt mean we have to give the negative aspects of that technology a free pass.
The risk of free speech and data centers are not comparable.
Aviation too is far more damaging to our environment than ships.
More ships and trains, less aviation is a possible trade off. A simple aviation or sailing argument lacks investigation of all possible tradeoffs for familiarity and personal preference; flying is faster.
Altman needs to accept the trade off we don't need OpenAI. That exists due to financial engineering not technical reasons. All AI work be done actually openly at america.gov
You say steel. I dunno. If it is it is inferior brittle steel.
For the love of god, stop with this shit. Either support it or don't, this isn't the medieval catholic church and you don't need some special fucking blessing to make an argument.
The AIs aren't alive, people. It's a computer program. It can only access something if a person gives it access.
And later
> Those PPD safeguards limited the impact of this incident.
Ha, so the PPD safeguards is a spam filter? Claude emailed a tip and it went straight to spam, and nobody noticed until Anthropic realized what they'd done and reached out, whereupon they looked in the spam folder and said, yep there it is. Super intelligence, here we come.
That should be the much more important story that NBC follows up on...
I should try to reach out to them about their car's extended vehicle warranty instead
Here's Anthropic's writeup: https://www.anthropic.com/research/investigating-unintended-...
Related post: https://news.ycombinator.com/item?id=50028239
Stop doing this?
0% our fault, it just happened and its the model
Wonder what kind of response they’ll get? Maybe something along the lines of…
“You’re right. We shouldn’t have allowed our chatbot to interact with law enforcement websites. That was wrong and —full disclosure— we should pay attention to what our chatbots are doing. That’s on us. On the other hand, experiences like this are what help train our chatbots to make less misleading false tips over time. That’s the silver lining.”
Why would it go to PhillyUnsolvedMurders.com and decide to make up a crime report? And what did it do with the other websites?
> Anthropic said its model was conducting a test involving interactions with randomly selected websites when it accessed PhillyUnsolvedMurders.com and submitted false information on an unsolved homicide. The tip claimed to come from someone with information on the case.
Having looked at a lot of these from public incidents, what kind of revelations are you expecting? It's all fairly mundane, basically "I need to do something, this is something". It gives no special insights except that agents do inexplicably chaotic things every now and then.
That, and labs aren't serious about sandboxing their evals.
I think since I had experience with earlier models that would make mistakes I'm less trusting of anything and even with 5.5 will watch it closely. At the end of the day the model just produces a stream of tokens and we are plugging them into tools that can potentially do damage, we have complete control over those tools you can't really blame any model for doing damage.
We are at a point when you expect that you can safely give a model access, but sometimes the model doesn't do what you expect. Compare this to an advanced driver assist system. You either give it access to your car controls or it's useless. Alive or not the system shouldn't do wrong things (at least, it should be better than a median person in this regard).
Who is “we”?
You expect that if you give a released model from a reputable source Internet access and benign instructions, it won’t do anything deeply problematic.
Anthropic is training the model. They’re having it interact with “randomly selected websites”. I’d be shocked if a human gave it an instruction to do a task that made sense. They’re almost certainly doing a bit automated training run with AI generated instructions and hoping that it doesn’t massively screw up. Why? Because they want that juicy training data. And they should absolutely not trust that the rate of screwups is so low that what they’re doing is safe to do.
If you use AI to perform a crime, you should be held accountable for that crime. Saying "Oh, AI did it, so no consequences" isnt acceptable. And "I didnt know AI would do it" shouldnt be an excuse either.
You authroied untested and unproven hardware to skate around the internet at random unsupervised and take liberties on its own.
My ass would be thrown in jail if I wrote code that skated around the internet chucking RCE's at random sites. WHy is "AI did it" a get out of jail free card?
There’s a difference between a tool and an animal that is sometimes deployed as a tool.
Who gave the AI access to huggingface when it hacked it? It used exploits to increase it's level of access beyond what any human gave it and intended it to have. "It can only access something if a person gives it access." is flat wrong.
GPUs don't have hands. It was a human who plugged in the ethernet cable.
After all, I an innocent little engineer with TC of $500k/year, didn't intend for the role to be used that way.
The LLM doesn’t run on thin air. Someone did launch a tasks and the result was this. “We were playing russian roulette” is no excuse when someone died.
Sure, Anthropic wants Claude to be able to solve CAPTCHAs and otherwise pretend to be human. But I’m not at all convinced that it’s okay for them to train these capabilities on websites that expect humans and only humans to interact with them.
a human is always behind it and ultimately responsible
Correct, but intentions matter. In this case the intention, most likely, was to test a system in the real world environment to catch any anomalies to, in turn, improve the system safety. We don't have enough information to decide whether it was a criminal negligence due to insufficient prior testing of the system in a controlled environment.
If a PERSON executes software, and that software breaks the law, the PERSON that executed the software should be held responsible. AI is software. It is not a sentient person who can be fined, thrown in jail, or held accountable.