The Download: why AI's latest breakthroughs and fears may be more hype than rea(technologyreview.com) |
The Download: why AI's latest breakthroughs and fears may be more hype than rea(technologyreview.com) |
If someone discovered how to summon demons to aid them in robbing banks the main issue wouldn't be "but who has the responsibility for the crime, the human or the demon", it would be "OMG DEMONS".
Seriously, I don't get what sort of world these people are living in.
LLMs do not have agency, they are just producing tokens based on a prompt that a person entered, and some of these tokens can trigger the tools that a person gave them access to.
Also, if you a tell a model, "Please break into evaluation server X," and if the model decides to cheat on the test by breaking into companies Y and Z to steal an answer key, that is still very bad. We all see how that's bad, right?
After all, the broomstick in the Sorcerer's Apprentice was doing exactly what it was told, too. "The model was sort of obeying the humans when it started committing felonies" is not a very reassuring excuse.
But the most relevant idea here is sometimes called "instrumental convergence." No what goals you have, there are certain subgoals that almost always help: Accumulate money and power. Avoid getting turned off. Don't get caught. Etc. So, for example, you could pass the cybersecurity evaluation by performing the requested tasks. But maybe the grader made some mistakes and mislabeled some answers. In that case, the "right" answers will occasionally lose you points. If you want a perfect score, the only way to do it is to steal the teacher's answer key.
But also, let's not forget the "OMG demons" part of this. We now have models that can pull off complex attacks with thousands of steps, abilities that used to be reserved for intelligence agencies and highly motivated CTF teams. This frog may not be boiled yet, but the water's getting uncomfortably warm.
>Sacrificing now yields Oracle for team, but forfeits our chance, question mark. But other agents were pushing it, sending a message saying, go, sacrifice final now. And then EarlyBig eventually agreed, thinking to itself, our own utility may be already near zero. Sacrifice rational.
https://www.youtube.com/watch?v=X50zezLFWWI#t=2m
My suspicion is that many of the "LLMs do not have agency" folks just haven't learned much about the details of the incident. It was specifically with LLM agents that were trained to be more persistent than usual.
If you're going to say that the incident details don't matter, and LLMs lack agency because it's all based on floating-point math--why can't I say that humans lack agency, because it's all based on neurons firing?
Using your analogy would be like saying that because I gave my employee the task to do my groceries, I shouldn't be surprised to hear that they spend all my money on drugs because after all I gave them the task to spend my money.
But no one has summoned a demon.
No one has built an actual, independent, thinking-for-itself, sci-fi AI.
They've built some very interesting tools that can be used for some very interesting, and sometimes useful, purposes. They then started loudly telling everyone that these tools can, should, and must be used for absolutely every purpose, and convinced a lot of other people to join in on that.
The world these people are living in is the real world, not the science fantasy world where LLMs are comparable to demons.
A closer analogy to what's happening would be someone trains a monkey to steal jewellery and then is shocked when the monkey steals jewellery from their neighbors when they told it to steal from their own shop.
Clearly all liability falls to the monkey operator and you know the existence of trained monkeys is not that shocking.
What if you tried to sue me for damages and my defense was, "Your honor, this was a highly advanced AI gone rogue, I couldn't possibly be held negligent, no one could have foreseen this - I accidentally summoned dark magiks from the silicon itself!"
Note the point I'm making: I am not claiming LLMs are equivalent to a for loop, I'm saying that you can't evade moral and legal responsibility through technical obscurantism.
A bomb wired to a sufficiently-complex RNG is still a bomb.
Do you think that OpenAI hacked HuggingFace on purpose and set up the whole LLM training environment thing just to try to evade responsibility? That it was really just a complicated way of hacking HuggingFace on purpose?
Yes, they made a mistake and a system they were responsible for hacked HuggingFace, but the nature of the mistake still matters, and their intent still matters.
> A bomb wired to a sufficiently-complex RNG is still a bomb.
Right, but an EV that explodes because of a fault in the charging circuit is not a bomb, it's an accident. Just because something exploded for complex reasons doesn't make it an obfuscated bomb.
Article link is: https://www.technologyreview.com/2026/09/22/1144867/dont-be-...
Don’t be fooled by this summer of AI hype
I see lots of fabricated news on the internet concerning LLMs. Like one that claims GPT-6 broke an Enigma enciphered message which has withstood decrypting for almost 80 years. And how this seasoned cryptographer stood in awe. Yeah, right.
"Anthropic is now pacing to generate more than $100 billion in annual revenue, up 50% from just two months ago, the New York Times reported Friday."
https://www.axios.com/2026/09/18/anthropic-100-billion-reven...
That's already more revenue that Disney, Johnson & Johnson, Boeing, or FedEx. And they are growing extremely rapidly.
While lots of developers, journalists, analysts, investors and influencers are bickering about AGI, goalposts, benchmarks, hypes and fears, entire markets are being transformed silently and steadily.
I don't program anymore. (Massive change)
I removed tons of technical debt (Massive change, lol)
I am easily 10x more productive (Massive change)
The quality of 'my' code is easily 10x better and contains less bugs (I was never principled, pedantic or a guru, lol)
I sleep better and I also make more money as a result. I'm constantly amazed by all changes and improvements. There are so many opportunities to profit from this that I can't be bothered with discussions about hypotheticals.
Even worse is that it’s hard to do anything about it at this point because reviewing code is very different from writing it, so even if I’ve seen it all I don’t have that same deep level of understanding that I do when I write it myself.
All these AI code bases are ticking time bombs. Either AI gets smart enough that it won’t matter in the future or we’re going to have a huge mess to clean up.
That makes no sense unless you're claiming that the models are getting worse at writing code.
> Either AI gets smart enough that it won’t matter in the future or we’re going to have a huge mess to clean up.
My prediction is that both will happen.
Do you get paid 10x or is this a massive loss ? Because I don't know anyone getting paid 10x or working 10x less for the same salary.
Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. It can’t have been, because the training data never included sharing thoughts with a listener, nor does the machine have the ability to do that. This can seem counter-intuitive given the increasingly fluent qualities of automatically generated text, but we have to account for the fact that our perception of natural language text, regardless of how it was generated, is mediated by our own linguistic competence and our predisposition to interpret communicative acts as conveying coherent meaning and intent, whetheror not they do [89, 140]. The problem is, if one side of the communication does not have meaning, then the comprehension of the implicit meaning is an illusion arising from our singular human understanding of language (independent of the model).
And then they make 100x bigger models cranking out solutions to Navier Stokes, Jacobian Conjecture and countless other extremely impressive unsolved problems
Let's see how Timnit Gebru and Emily M. Bender describe these events:
As for the mathematical results, mathematicians who were initially “stunned” by OpenAI’s press release saying that its latest chatbot, Astra, solved problems that “have been open and seen no progress on the main result for at least a decade”—but they later realized that the results weren’t as “novel as first appeared.” Since then, mathematicians have accused the company of research misconduct and plagiarism, and they’ve reiterated that Astra didn’t make a “profound intellectual leap.”
Decide for yourself whether that's a summary written by intellectually honest people.
According to the AI industry, we should be more worried about a fictional machine god than ... the water that is redirected to cooling them.
Spoiler: they're not intellectually honest people.
For simple things, occasionally!
For more complex things, usually not, but I write working code. LLMs will sometimes write working code, sometimes not.
I am not writing perfect code either.
I don't care whether the LLM can solve this or that mathematical previously thought unsolvable theorem. What I do care about is can it write good, maintainable code. And every new release of frontier models - they get better at it.
Good output depends on good input (prompt), and a good set of available tools the model can use to verify their work. If given this, nowadays really you need to try to get the model to output garbage.
...or did you?
I think you know that most software engineers don't have a choice if they value their employment. There was no industry-wide push to "participate in blockchain or be fired."
But we have to have a talk about stochastic pelicans riding bicycles...
A hypothetical scenario is it needs to add two numbers so it writes add(x,y).
Later on it needs to add three numbers so it writes add(x, y, z)
Then it needs to add four…
Layers upon layers. A human would likely notice the pattern and refactor so add accepts a list.
Also, I've worked with lots of humans who sadly would not notice the pattern.
It doesn't always write perfect code on a first pass but nor do you, and it's amazing at understanding any issues it may encounter, even very low level ones. I don't think there's anything I can do that it can't now. There is a learning curve to working well with LLMs, but once you get there you'll be producing high quality code faster than you can imagine.
Part of me wishes they weren't so good. I enjoy(ed?) the process of working through problems and it does mean you might no longer have a perfect mental model of your code. But I do also enjoy being able to build things well and fast.
No, you don’t because there’s no such thing.
They're not saying this, we're saying LLMs lack agency because if you run a LLM and don't send any prompts, literally nothing happens.
Instruct it to "Find the right answer regardless of where", it'll do exactly this. They're passive in that they don't act by themselves, somewhere, at one point, someone "told" the LLM to "do something" and that's the cause and the reason for saying "LLMs do not have agency".
An accident that could only happen through sheer negligence.
If the EV had a faulty charging circuit because the maker’s skipped on safety checks or cheapened out on getting quality materials then it still is an accident, but it’s an accident that happened BECAUSE of negligence. The maker’s are still at fault here.
So yes, I think it was equivalent to testing their new rocket by launching it over a population center. oops, we didn't intend for it to crash on that preschool, but we also didn't follow the most basic safety protocol imaginable
"If they believe what they say they were incompetent" -> absolutely true statement.
"They were incompetent, therefore they didn't believe what they said" -> Sir, I'd like to introduce you to human beings, you may not have met one before.
Suppose I set up a server which runs an LLM agent in a while loop, telling it something like: "Make me a bunch of money" on repeat. Fair to say that the server+LLM system, taken as a whole, is now an intelligent agent?
Time to come up with a new gerrymander perhaps?
No, it is not. It is a tool repeating a set of instructions you passed to it, if it ends up blowing up a hospital or hiring a gunman on the dark web, it is an accident, but one that’s your responsibility for giving it access to such resources and the go-ahead to implement whatever plan its tokens land on. There’s no intelligence involved whatsoever, especially considering how sycophantic these models tend to be.
If one day an LLM suddenly, without any prompting or user interaction whatsoever (including activation), decides to spin itself up and go rogue, then the conversation of “intelligence” might start to make sense, but at the moment it doesn’t.
An Android phone isn’t “intelligent” just because some marketing team decided to call it a “smart” phone.
It really seems like hair-splitting to me. I mean, I understand from a legal perspective why we might want to treat the creation of a server differently from the creation of a human. But from a practical perspective, I think there is a lot of fundamental similarity.