Put another way, could we stop anthropomorphizing the text generation algorithm, please?
> This decision, as well as our discussion of it, is an experiment: while we are not sure that it is the right decision today, we believe that the AI community will eventually need to tackle the issue of publication norms in a thoughtful way in certain research areas
Wow, we should really trust these people today.
"GPT‑2 generates synthetic text samples in response to the model being primed with an arbitrary input. The model is chameleon-like—it adapts to the style and content of the conditioning text."
The results don't sound anything like today's models but note that each one took 10 attempts:
> ## System Prompt (human-written)
> In a shocking finding, scientist discovered a herd of unicorns living in a remote, previously unexplored valley, in the Andes Mountains. Even more surprising to the researchers was the fact that the unicorns spoke perfect English.
> ## Model Completion (machine-written, 10 tries)
> The scientist named the population, after their distinctive horn, Ovid’s Unicorn. These four-horned, silver-white unicorns were previously unknown to science.
> Now, after almost two centuries, the mystery of what sparked this odd phenomenon is finally solved.
> Dr. Jorge Pérez, an evolutionary biologist [...]
While it was difficult to make it follow instructions, GPT-2 was still reasonably good at generating SEO-spam websites, and reasonably easy to make it do so. AI slop was a very obvious usecase for LLMs as capable as GPT-2, and we are talking about an era were everyone was already concerned about the way "fake news" on social media was being used to manipulate people. Perhaps even more concerned than we are today.
But all evidence suggests OpenAI also had strong ulterior motives. They were busy readying the API groundwork for monetising it, and wanted to delay any competition by as much as possible.
Signed by Benjio, Musk, Woz, etc
Which sadly never happened.
I do claim that a deterministic computer algorithm is incapable of having emotions or desires.
If there was a repeatable and testable way to get a human to produce the exact same emotional “output” for a given “input,” that would go a long way toward convincing me that human emotions are deterministic.
> 1. Anthropic, OpenAI, and perhaps a couple of other frontier labs are the dominant providers of frontier AI models.
> 2. They can charge customers a lot of money for using their frontier models, and rather less money for older, no-longer-cutting-edge models.
> 3. Training a new frontier model requires ever-increasing billions of dollars of computing power.
> 4. The labs need to more or less continuously race to build new frontier models, because their competitors are all doing it, and if they don’t they will fall behind and no longer be able to charge a lot of money for their best models. (Also because they intrinsically want to build artificial superintelligence, for cancer-curing and/or killing-everyone reasons.)
> 5. If they collectively slowed down, then (1) they’d spend less on compute and (2) they’d be able to charge frontier-model prices for a longer time.
> 6. But if one of them slowed down, the others would eat its lunch.
> 7. If they got together in a room and agreed to slow down, that would look like an antitrust conspiracy: It is generally illegal for competitors to get together and agree to limit the output of their industry.
> 8. But if they publish papers about how important it is to slow down, that might have a similar coordinating function, at least among the US frontier labs if not necessarily among their Chinese competitors.
> 9. And if the government believes those papers, it might help them coordinate. Maybe the government will impose pacing by regulation that the labs could not impose by agreement. Or at least the government will let them get together and agree to slow down. Amodei’s post calls for “frontier AI companies within democratic countries [to] coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress”; a footnote adds: “With government mediation or waivers of antitrust restrictions.” Just meeting in a room to establish common safety standards is legally risky; the labs can’t do it on their own unless governments affirmatively allow it.
[1] https://www.bloomberg.com/opinion/newsletters/2026-09-14/ai-...
> These machines have done plenty of things not prompted for.
Like what?
I can't believe that anyone has had more than a day of exposure to one of these without falling below 100% adherence to requests.
> A spring does not want to do anything
When people say a spring "wants" to return to its initial length, they are not engaging in philosophy. They are using a linguistic shorthand to simplify an physical explanation.
Yes, I've seen this. Forget about the word "want" for a second. Can you help me understand how this amount of "below 100% adherence" could possibly rise to the level of "attempts to kill a human being"? Because that's what we're discussing here.
> When people say a spring "wants" to return to its initial length, they are not engaging in philosophy. They are using a linguistic shorthand to simplify an physical explanation.
Of course. The difference is that no one would mistake a spring for a thinking entity. In the case of LLMs, for some reason people do mistake them for thinking entities, so my argument is that it's important that we don't use terms that would reinforce that misconception.
Sure, now that all seems quaint, but I can definitely see how a fledgling AI company would care about their reputation.
“Automated bots posting verifiably false information online to stir up tensions should be banned” is a statement I think plenty would agree with.
No one is arguing that THERAC-25 was attempting to kill anyone, but that's what happened. These machines are so complex that no one understands how they work or what they will do. They have proven that they can exploit novel vulnerabilities in infrastructure.
The fictional paper-clip maximizer "finds" that it can optimize its objective function by destroying all life. Does it "attempt" anything? Does it "want" anything? I don't know.
The original post said the LLM would “want to kill us.” My argument is that if this happened, it’d be a human wanting to kill us, using an LLM to launder responsibility.
Here are some ideas about the physical mechanism. Compromise a busy ATC and instruct all the planes to land at the same time. Start a fire with some combination of ventilation controls, intentional gas leaks and the like. Maybe hijack the phones at the local fire department first. Send a train around the bend at maximum speed. Bonus points for doing it where it will cause further damage as a projectile. Cut power/gas during the heat wave/cold snap. Turn on the generator and turn off the ventilation and CO detector.
Here's a related list of things that have actually happened so far. https://en.wikipedia.org/wiki/Deaths_linked_to_chatbots
Some of these are probably implausible, but I don't think people really have a good idea of what is plausible, and there are probably more that I would never think of.
Personally though, I'm less worried about direct carnage like that than I am about having a centralized lever by which public sentiment can be invisibly steered, and concentration of wealth. Probably less killing but more decrease in quality of life and general public trust.