What Emily Bender meant by "stochastic parrots"(spectrum.ieee.org) |
What Emily Bender meant by "stochastic parrots"(spectrum.ieee.org) |
Not suggesting that I don't say stuff on autopilot sometimes but for many people, it's their only mode of operation. They never actually think about anything from first principles. Their whole approach to language is just chaining catchphrases together. It's how a toddler thinks; it seems like many people never moved past that stage of development.
- Lots of Haiku around, many mistakes unless process is very clear - Some Sonnets, still do mistakes but can adapt - Some Opus, able to improvise and think outside the box.
But even the Human Opus/Mythos are hilariously wrong sometimes.
There's a lot more happening behind the scenes when a human repeats phrases than what's happening in an LLM.
Sociological phenomenon. The desire to be liked, successful, or popular. The feeling that those phrases brings up.
LLMs are not experiencing any of that. As far as we know, neither is a parrot.
Did I not, then, as I grew out of infancy, come next to boyhood, or rather did it not come to me and succeed my infancy? My infancy did not go away (for where would it go?). It was simply no longer present; and I was no longer an infant who could not speak, but now a chattering boy. I remember this, and I have since observed how I learned to speak. My elders did not teach me words by rote, as they taught me my letters afterward. But I myself, when I was unable to communicate all I wished to say to whomever I wished by means of whimperings and grunts and various gestures of my limbs (which I used to reinforce my demands), I myself repeated the sounds already stored in my memory by the mind which thou, O my God, hadst given me. When they called some thing by name and pointed it out while they spoke, I saw it and realized that the thing they wished to indicate was called by the name they then uttered. And what they meant was made plain by the gestures of their bodies, by a kind of natural language, common to all nations, which expresses itself through changes of countenance, glances of the eye, gestures and intonations which indicate a disposition and attitude--either to seek or to possess, to reject or to avoid. So it was that by frequently hearing words, in different phrases, I gradually identified the objects which the words stood for and, having formed my mouth to repeat these signs, I was thereby able to express my will. Thus I exchanged with those about me the verbal signs by which we express our wishes and advanced deeper into the stormy fellowship of human life, depending all the while upon the authority of my parents and the behest of my elders.
[https://faculty.georgetown.edu/jod/augustine/conf.pdf]Humans learn language opportunistically. Toddlers start with a powerful "superchimpanzee" understanding of the real world, and use that to learn words in order to satisfy their needs and desires. Statistical frequency is incidental to what words a toddler learns: what matters is the real-world context. Also note how important it is that infants instinctively understand nonverbal communication.
The most depressing thing about the 2020s AI summer is watching ignorant tech workers use the success of LLMs to launder their own ignorant misanthropy. Your views are many many centuries out of date.
> "The basic topic of my address today concerns how much of cognition is in the head of the infant and how much in the mind of the theoretician. My general stance is that we are being treated to an interpretive flavor of infant behavior that is much too rich."
—— Who put the cog in infant cognition? Is rich interpretation too costly? https://home.fau.edu/lewkowic/web/Haith_Critique%20of%20Cogn...
> Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot.
I think this metaphor is so strained as to not be useful. I think key here is that the authors say "without any reference to meaning", which is a heavily loaded term, that does definitely apply to parrots, but does not apply when you apply it to immense bodies of text.
Namely that language embeds meaning in language. A sentence being written by a human (as a starting point) is designed to have consistent meaning. While it is possible to write syntactically correct meaningless text, that is not what most of human language has done; the meaning cannot be removed from the text.
This I think is clarifying, from the same paragraph in the text:
> ... the training data never included sharing thoughts with a listener, nor does the machine have the ability to do that.
That's just facially incorrect. The training data is entirely about sharing thoughts with a listener. Else why is the text being written?
> Meanwhile, O, a hyper-intelligent deep-sea octopus who is unable to visit or observe the two islands, discovers a way to tap into the underwater cable and listen in on A and B’s conversations. O knows nothing about English initially, but is very good at detecting statistical patterns. Over time, O learns to predict with great accuracy how B will respond to each of A’s utterances. O also observes that certain words tend to occur in similar contexts, and perhaps learns to generalize across lexical patterns by hypothesizing that they can be used somewhat interchangeably. Nonetheless, Ohas never observed these objects, and thus would not be able to pick out the referent of a word when presented with a set of (physical) alternatives.
This seems kind of obviously wrong at least in the context of coding agents. These models get trained on actual output of the previous version of the model doing its job, often "IRL" on a real computer/project. It's like O is in the conversation for years now and learning from his own interactions between A <-> O <-> B, where A is the human and B is the computer.
The idea O ontologically has never "observed" "these objects" or referents is philosophically strained. Have I observed the moon, or a finger pointing at the moon? Have I observed `sed` more than Fable?
I don't understand what we're setting the record straight on. This is the core point of dispute, and the author just blazes past it to focus on other things. I'm glad to hear "stochastic parrot" isn't intended as an insult, and I agree that it's not right to think of LLMs as a box with a little homunculus inside replying to you. But to me it seems obvious that LLMs are not repeating patterns without comprehension and do understand what they are saying; otherwise they would not be capable of doing things they routinely do.
Cognitive skills such as tool use and complex navigation predate language as well. That means there's a core of reasoning in humans that doesn't depend on "tokens" or "language" of any kind. Language is a tool for communication and forming complex human societies, but it's not cognition.
> The thing is a parrot just says things it has already heard, it doesn't perform complex reasoning on novel situations and then explain it succinctly.
Well a parrot does perform complex reasoning on novel situations all the time. It just doesn't have the wiring to connect that to "tokenized" human language. I suspect LLMs have the opposite problem, where they exist in the domain of their "tokens" and have no way to connect these to truly novel situations that have no existing words to describe them.
There are people who argue semantics, that we can call the pattern matching that LLMs do “understanding”, or the moronic “how do we know that’s isn’t all we do” but for the normal use of comprehension, LLMs at a fundamental level don’t.
there is an entire genre of riddles based on this kind of misdirection that works on humans, such as "A plane crashed on the border or US and Canada. Where do they bury the survivors?"
As with humans, comprehension is not a binary property of the agent - it is a quality that can be present in some situations and absent in others. LLMs may emit correct outputs sometimes because they do comprehend the input, and emit incorrect outputs in other cases when they do not comprehend the input.
In order to show that LLMs can't comprehend, we'd have to show that there are no (or at least very few) situations in which they exhibit comprehension, not show that there are some situations in which they don't.
So this seems obvious to you, and yet to many others, it is equally obvious that LLMs can/could do the things they routinely do without any meaningful sense of "understanding".
It's not like you can be agnostic, or measured about this. It's like someone explaining a car to you, saying, "look here is where you put the fuel, here is where it ignites, where the axels are turned..." And you, trying to be measured, are like "hm well yes of course that all is clearly important, but there is clearly just a bit of magic here somewhere, between all the different 'parts'."
Is it possible you're making the following error described in the article?
> The fact that these systems are designed to mimic the way we use language makes it very easy for people to mistake them for other people.
Clearly you don't believe it's actually a person ("it's not right to think of LLMs as a box with a little homunculus inside replying to you"), but you do believe it's doing something a little bit magical. Is it possible because the interface is linguistic, and every other thing in your world that communicates with language is intelligent, that you're projecting something that just isn't there onto the situation?
I'm sorry if this line of questioning is a little invasive. But this is literally the "danger" the original paper talks about, and it seems an awful lot like you've fallen for it.
> What would happen if I walked to the top of a skyscraper with a soda can full of Maraschino cherries and let them go?
And its answer (https://chatgpt.com/s/t_6a4bd9ffa5708191901bb6d43c89f43b) clearly demonstrates understanding. It knew that this is a dangerous thing I should not do in real life, and that my question is ambiguous about whether I intend to drop the can, and that this might be intended as a physics problem rather than a real life scenario.
Where can I access such a Lisp expert system?
If I cannot because they don't exist: then they cannot do the same things an LLM can do. And of course one can assert anything and everything about what a non-existing thing could do.
...
> It knew that this is a dangerous thing I should not do in real life
From the ChatGPT response you linked, all I see for sure is some matches on the following patterns:
drop $thing from skyscraper --> bad behavior
drop $thing --> physics
can of $stuff --> contents in/out of can
Then there are some sentences of likely characters following those patterns. You don't need anything more than a basic cartoon-level understanding of how an LLM works to explain this output. I see no evidence of reasoning or understanding here, or any theory of "real life".It also does an incredibly poor job of answering your question. It makes no attempt to explain what might actually happen. If it has been trained on the entire corpus of medical science, and it is indeed intelligent, then surely it can reference ballistics studies and give you a very detailed and thorough theory of what--exactly--injuries you might expect from a 12oz can being dropped from the height of a skyscraper. Calculating the terminal velocity and therefore the momentum of the can is trivial. Characterizing the physics of the impact on various parts of a human body is trivial. If it actually understood your request why didn't it just answer the question?
It's a rhetorical question. LLMs do not "understand". It is completely outside their capability. "Understanding" is something we impose upon their output (to loosely quote TFA). [edit] I think the most powerful evidence for a lack of any understanding whatsoever is all the stuff about the cherries being in or out of the can. Yes, cans contain things. That is not a profound observation, nor is it at all relevant to the question. If you drop an empty can off a skyscraper nothing meaningful will happen. And, no, probably dumping all the cherries out won't hurt anyone or cause a slipping hazard... It's also not particularly relevant to point out that dropping things off skyscrapers is bad behavior. But that's more forgivable from a CYA standpoint.
I believe you are projecting something that is not there onto a completely mindless stochastic process.
[another edit] I found this helpful, you may also: https://www.0xkato.xyz/how-llms-actually-work/
I think LLMs understand without reasoning. They've built a large associative network of concepts (a kind of understanding), but we don't yet have a good handle on the process of reasoning using that network.
The child knows how to catch the ball, without understanding. Later, the child learns both reason and physics, and can reason about ball catching in a different way.
I don't think that it is useful to say that LLMs understand anything they say, or that we say to them.
I disagree that decoupling "knowing" from "understanding" is well motivated or useful in any way.
> I don't think that it is useful to say that LLMs understand anything they say, or that we say to them.
This is starting to seem like a Chinese Room disagreement, so I'll just say that what that thought experiment demonstrated is that there is no robust or meaningful way to distinguish semantics from syntactic processing given everything we know about reality, and that LLMs were a fantastic demonstration of this illusion, eg. that "semantics" are reducible to a relational network of syntactic elements.
I remember being angry about this situation when I first saw it on social media, until I read the details: This person submitted a list of demands to her employer and said that if they weren’t met, she quit. Google wasn’t going to meet her demands so they considered it acceptance of her resignation. There has been a movement trying to debate whether it was a firing or resignation ever since.
The original paper they published gets recirculated every year or two as some landmark history of AI safety, but as other commenters have noted it wasn’t really a great paper nor was it groundbreaking at the time. If not for the controversy surrounding the resignation/firing (depending on your POV), I don’t think it would have been notable.
I think she's since since lost a lot of her allure, especially when she didn't change her mind when the facts about the AI water usage changed 1000x
Did she really think Google cares about ethics? Such positions seem purely performative, we all know that ethics go out the window first to make room for more profits.
She was part of the "Ethical Artificial Intelligence Team" of what was then, and still is now, one of the corporations World wide spending the largest amount of resources precisely on using AI commercially.
Meanwhile, the paper has 2 points of criticism towards AI. 1 is a bunch of carbon consumption complaints assuming NVIDIA cards with coal-fired power, while a lot of effort at contemporary Google went towards getting TPUs running on green power. I suspect this was what people wanted to object to, a lot of effort went into those green power projects and she was just denying it. The complaint seems prophetic now but it was not true about Google then.
The other criticism was about which language the LLMs use, they average the input data of normal humans instead of talking the way the paper author thinks they should talk. The phrase "women doctors" is called out as problematic. I'm less inclined to think people objected strongly to this given the zeitgeist at the time, it was probably people who worked on the green energy projects and were pissed off that their contributions were ignored, but still, nobody elected her Queen of English, she can have her opinions but she's not a victim for not having them adopted by everyone.
(BTW, quite bold to say input data from Reddit and 4Chan is how “normal” people speak. There is a lot of language in the training data of any model you really do not wish your application to use ever.)
Google forced its researchers to retract an already submitted paper because it undermined its strategic and commercial story around large language models. The "we just accepted her resignation" is just a lie. Google made harsh demands with opaque reviewers that made vague objections, and then Jeff Dean moved very quickly to get rid of Gebru. Other Google researchers reported that they usually got to work through objections, Gebru got no such opportunity. Google showed that AI labs will not tolerate internal research that seriously criticizes technology central to its business.
Apparently google did not even give them the freedom to come to their own conclusions.
I just skimmed through the paper again and what it says seems to hold up well.
The paper is heritical in that it suggests slowing down, using carefully chosen less biased data, and understanding how it all works.
Basically challenging the AI bitter lesson.
I can see how that would meet with a lot of resistance from people who have gone all in on the bitter lesson.
Gebru is not a researcher. She is a modern-age Trofim Lysenko, who politicizes everything and wields political correctness as a weapon to purge any dissent.
I'd say what's under debate is whether uncritical LLM adoption is mainly unethical or mainly religious.
>Timnit responded with an email requiring that a number of conditions be met in order for her to continue working at Google, including revealing the identities of every person who Megan and I had spoken to and consulted as part of the review of the paper and the exact feedback. Timnit wrote that if we didn’t meet these demands, she would leave Google and work on an end date. We accept and respect her decision to resign from Google.
This is Google's side of it; I think the following is a fair piece of primary-source journalism if you want to go deeper:
https://www.platformer.news/the-withering-email-that-got-an-...
I have watched it happen multiple times that someone from academia joins a research group at a large corporation and finds out to their chagrin that they can't just overtly bully colleagues as easily without tenure.
edit: And it looks like Mitchell immediately left an angry comment about being called a "linguist" despite the article never doing so. Starting to notice a pattern of confrontation, ego, and escalation among some of these people.
I appeciate short letters like this that get straight to the point...
On a completely tangential sidenote, octopusses are actually very very intelligent: https://www.nhm.ac.uk/discover/octopuses-keep-surprising-us-...
Almost all other particularly intelligent animals seem to be gregarious, and it's easy to conclude that a social lifestyle tends to select for more intelligence, a sophisticated theory of mind, and so on (I like to think that that's exactly what was responsible for a runaway intelligence explosion in humans). But in the case of cephalopods, there's something else that has been applying selection pressure towards exceptional intelligence.
OpenAI offered ChatGPT to the world. A large, monied cross-section of the world had yet to throw its capital behind the Large Language Model technology that made the ChatBot possible. While it is fair to see AI development now as a global imposition, OpenAI did not have the agency as a 2022 startup to impose on the scale we see now.
I asked Mistral, and it guestimated that Altman, Thiel, Musk, and Hoffman had like $20.3B together when they founded it. Sound to me that the founding of OpenAI was exactly the point when the monied world threw its dollars behind AI.
I think this paper would have been best split off from the conjoined criticism of environmental effects (which could have been its own paper, but not one published by Google, since their leadership's fundamental beliefs disagree with the paper's environmental impact premise. And the remaining part on text models could have been a bit more focused on the technical issues associated with statistical text processing and meaning, rather than criticism of the power structure that is loosely associated with the current AI push.
The authors of the paper make some good points (in it and elsewhere), but I've seen precisely nothing that suggests they aren't awful people who resort to playing the victim card, inciting cancel mobs and in general behaving like the worst stereotypes of "leftist academia" in response to any criticism.
"Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot."
This was not even a correct criticism in 2021. She is right that, at the time, the pretraining -- where it learns to predict missing words in pre-existing corpuses of text -- is basically a stochastic parrot.
But nowhere in her paper does the term "reinforcement" come up. At the time, this was done mainly through RLHF (reinforcement learning from human feedback) - after the initial training is done, you then tune the model's responses based on human grading. Humans imbue their own meanings into the parameter weights through their judgment.
At this point, they aren't really stochastic parrots anymore, because parameter weights have been shaped beyond the text corpus. It's not purely probabilistic in the sense of using the probabilities of the underlying text sequences. (It still is probabilistic in its output, but that is a pointless claim, because all events in the universe are also probabilistic; it is not enough to merely claim that probability is involved in some way in the outputs.)
RLHF was already in use prior to the paper, and was written about by Christiano in 2017 "Deep reinforcement learning from human preferences," so it's surprising that Bender apparently didn't know about this well-known paper.
RLHF was also, of course, a precursor to a more advanced form of parameter shaping - reinforcement learning with verified rewards, or RLVF, which has driven a lot of the gains in verifiable domains lately. That was not done in 2021 when she wrote the paper. But if you knew about RLHF -- and knew how Alpha Zero worked, with training neural nets on game rollouts -- you could squint and see that it might be useful for language models.
So after being proven to not only having a limited understanding of the field at the time, but also not being able to forecast the field, she's now walking back what she meant by "stochastic parrot," I assume because she believes readers will not read what she wrote. But despite the protests, her original claim was that it is a parrot because the text has no meaning -- a direct quote from the paper, which only really makes sense if training stops at the pretrain.
In her previous interviews, I've found her assertion that LLMs aren't useful and will never be good at anything totally uncompelling. Also laughed at this quote as she's been pretty harsh IMO on "the people who like the systems".
> it’s all about trying to make vivid to people who aren’t in the business of building language technology what these systems actually do, which is not the same thing as insulting the systems or insulting the people who like the systems.
At a minimum it’s probably more accurate than “AI”.
At this point, I think the authors are really counting on people not to read their paper.
You can't. That's just the way the news goes.
Essentially, what you are saying is that because some people somewhere frame a statement as pejorative, the statement itself is inherently pejorative. By that logic, every criticism ever articulated is inherently an insult.
Bailey: well your version
Which one it doesn't do for you? Does it sometimes answer 4, sometimes 5?
There are definitely models that will always give 100% of the time the exact same answer, bit-for-bit, given the same input and seed. There are generative image models you can run locally doing just that. But you can also run some the SOTA chinese LLMs at "temperature 0" and, given the same input, they'll always give you the exact same output.
Because it's just a machine doing computation.
In the beginning of LLMs some "engineers" have tried to hand-wave non-sensical explanation as to why LLMs couldn't possibly be deterministic but: the open-weights models that can be run in a 100% deterministic way are way more powerful than the SOTA models of back then, so those explanation were pure rubbish bollocks.
Now of course if you run a complex chain of events, with LLMs doing calls to other LLMs, where some of them go fetch infos on unreliable networks, with infos that may have changed, then, logically, you won't always get the same answer.
2. LLMs are detemrinistic. They have a parameter to tune how stochastic they are.
So stochastic parrots could indeed be a good description of LLMs. But I think that she meant it as a diminishing term (against the technology) which is pointless. Probably more of a reaction against SV tech bros than more nuanced interpretations.
“Stochastic parrot got picked up and interpreted by other people as a minimization or an insult. It was not meant that way. Other people might be using it that way but that’s not how I intended it”.
Yeah that’s because it was chosen to be an insulting phrase.. Parroting is only ever used as a pejorative phrase. But sure, everyone else mindlessly parroting this line is the problem here.
This paper was always lousy, but it has really not aged well. We are living in a world when where an LLM has solved an Erdos problem. In a world where LLMs produce novel results that rival human thinking any conceptual reduction of an LLM is going to start inviting some unpleasant comparisons with human thinking.
There have been a few years now of arguments about the level to which transformers do or do not have a world model (v.s. being purely stochastic parrots like early pre-trained LLMs) and now we have some tools to actually make quantifiable determinations.
I think "artificial" is actually a pretty good term to describe the output of the models. That output does appear to resemble at least some definition of the word "intelligence" - there is some ability there to do cognition over information that's been provided to them in-context.
What is it to understand, then? If they can work in complex domains and produce coherent output, it would seem to necessitate at least some definition of "understanding" of the corpus, even if that understanding is unlike how a human's brain would understand it.
What else should we call them then? They model language and information in ways that allow them to manipulate it on the fly. They do so 'unnaturally' from a human's point of reference.
I legitimately can't come up with a better term than 'artifical intelligence' -- not to be confused with artificial consciousness, which I don't think exists (yet).
Now if they had said, "Imagine your average American ..." (/s)
It is true that LLMs often behave haphazardly, and do rely on statistics. But plenty of research has shown them behaving in methodical ways too. There are findings going both ways!
Granted, many of the strongest contradictory results appeared after the Stochastic Parrots paper, so it isn't like they were ignoring the literature at the time. But they did make a very strong claim, and in the half-decade since, a lot of evidence has come out against it.
We’ve know that since 1943 when McCulloch-Pitts came up with the first “artificial neuron” definition. And since LLMs are a descendant technology — our assumption should be they’re reasoning in some internal learned logic.
This is what the evidence supports — eg, the “stochastic parrot” crowd never can explain transfer learning. Whereas for the internal reasoning crowd that is easy: removing your top level judgments from a theory still leaves you with useful terms for describing a new theory — eg, removing your judgments about “which animal is this?” but preserving the underlying structure for representing an image in your new judgments, “is this cancer?”
There’s 80 years of reason to think DNNs reason and zero support other than “sTaTs R mAgIc!” to support the stochastic parrot interpretation.
Ignorance isn’t argument.
They never are. Ever.
And even when they are: they sure seem to bet against Moore's Law or just the general tendency for things to get better/efficient over time.
It's frankly remarkable how capable the models have become that we can run locally now on a decent laptop.
The same thing happened with image generation. I've had arguments with people that image generators are killing the environment, but I can do it in 20-30 seconds on my GPU. No one bats an eyelash when I play 20-30 minutes or even hours of a video game on my GPU, but the images are burning down the planet.
It's slightly maddening.
It's not a criticism of the paper itself, but multimodal models came shortly after and provide grounding that is more of the sort the paper is getting at, and it didn't seem like anybody updated on that at all. If multimodal models were still stochastic parrots by the original argument, humans would have to be as well; we don't have any way to ground anything beneath sense data and evolution can't have programmed some innate grounding into us because it didn't either. But (and maybe this is my own misperception) nobody threw in the towel at that point.
I confess I never read the original paper until now, opting to absorb by osmosis instead, and I was quite surprised that they don't really make a deeper case than that. After just a few paragraphs about how they can't be grounded because humans don't express their thoughts directly, it lurches into a page about how they can be biased by training. And they certainly can be, but that has little to say about their stochastic nature- humans are biased as a rule with no exception. (For the record, I only read the Stochastic Parrots section before this reply.)
It's not really a bad paper, but I don't see why it ever carried the esteem it did. Hating on it is like hating on Taylor Swift- she's fine, yes, but for her level of success, one is inclined to question every dumb lyric where others get a pass. (Apologies to Swift fans, substitute a successful artist you don't care for here.)
did you think this through?
imagine the sentence was "This sentence has four words", now extrapolate that to all the shit that can exist in a dataset and train a model on that dataset - do you know what will happen? - go ahead and think it through.
> LMs are not performing natural language understanding (NLU), and only have success in tasks that can be approached by manipulating linguistic form
... which is presented as unarguable fact, yet is untrue. It was obviously wrong at the time it was written and it's been proven wrong in many ways since. Worse is that they're still at it. In the article she's saying:
> Q: What are the most common misconceptions about the “stochastic parrots” metaphor? Bender: I think one of the biggest ones is, “Bender says AI is a stochastic parrot.”
Her name is on a paper titled "On the danger of stochastic parrots". It has a section titled "Stochastic parrots" and in section 6.1 it says:
> An LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot.
She did say that, as clear as day. Now she's trying to rewrite history. Ugly behavior.
If you train it on lots of working code, then it's useful for coding. If you trained it primarily on non-working code it would produce nonsense.
>> > LMs are not performing natural language understanding (NLU), and only have success in tasks that can be approached by manipulating linguistic form
did you read the paper carefully??? this line is directly cited from another paper (pay attention - it's from july 2020) . the full line is - LMs are not performing natural language understanding (NLU), and only have success in tasks that can be approached by manipulating linguistic form [ 14 ]
the paper its cited from is titled - Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data
>>>She did say that, as clear as day. Now she's trying to rewrite history. Ugly behavior.
did you read the article after that line or was there a shortage of attention span??
And secondarily, and maybe only partially the authors’ fault, is the enormous tidal wave of morons that this paper minted who plague us with their misunderstandings to this day.
>>And my criticism is that if you observe Claude code with opus 4.8 working through an entirely novel problem that nobody has ever worked on before and which certainly wasn’t in its training data, the choice to even metaphorically call them stochastic parrots turned out to be egregiously wrong.
First of all, for your own benefit - Claude code stopped showing the real reasoning. It only shows a summarized version of it now. So don't ever ask someone to observe the model. Second, do you know when this paper came out? do you know when opus 4.8 came out??? how do you know what is novel? did opus 4.8 tell you it was novel? how do you know no one has worked on it before?
Agreed, it's a problematic term that conflates theoretical research with better search results.
Deep learning and machine learning (ML) are both unromanticised (un-hyped) terms.
In some cases, it's just tooling with a better interface. As we have done with other complex computing systems (e.g. Deep Blue, Watson), we might just end up naming it a computing system for querying, e.g.
https://en.wikipedia.org/wiki/LCARS
We would like a neutral term for such a system, and in this sense it's better to call it "A.I." than to make a verb from google or bing or other corporate name.
But in refined cases, such tooling may offer a credible (as in "believable") human experience, like Weizenbaum's Eliza (https://en.wikipedia.org/wiki/ELIZA). Some users of Eliza fully believed that Eliza listened and understood at a profound human level. Eliza was a simple computer programme.
Computing systems are not humans. They have no accountability in real life. People may be so comfortable with the user-experience that they cannot distinguish it from interacting with another human. That doesn't make the computer system alive and accountable. And if it's run by a corporation, it will almost certainly make big promises while energetically seeking to avoid accountability. ^_^
Source?
This question depends on how you define research productivity. There is close to two hundred AI papers published every weekday. Most of them are about GenAI. Most don't seem to be all thay good. The progress in actual model improvement had mostly stalled. If you interact with the latest "raw" models they display all of the issues we've seen in GPT-3.5, just at a smaller rate. The "amazing gamechanger breakthroughs" I read about on social media every week do not seem to lead anywhere. It's all kind of boring, really.
The new "hotness" in AI is clearly building more and more elaborate harnesses. This is not at all the direction AI boosters have predicted couple years ago.
Personally, I think the "stochastic parrot" mental model is far more useful for science, because it primes people for proper testing, skepticism and researching alternatives. If you want useful AI, you want people working on it being skeptical, not credulous.
To me the real question begins only once we have a clear example of a non-trivial scientific discovery that is implicit (IE, not an obvious outcome of reading the literature and talking to the experts) and experimentally verifiable. Once that happens- especially if it is a reproducible process (IE, more discoveries) and it's significant (IE, impacts human life and mind in some profound way)- then the onus very much lies on Bender and her coauthors to explain whether we need more than a sufficiently advanced stochastic parrot.
By that standard, parrots, and it's not even close. The framing of intelligence led to an enormous number of predictions that simply haven't been realised: an end to all white collar work, UBI, a total revolution in society, a literal robot god.
People are so desperate to view 'stochastic parrots' as dismissive that they misread the original argument while quickly ignoring all the failed predictions about how AI was going to overturn, save, and destroy everything.
Once in frustration I called a certain frontier model "Sam Altman's Tin Bird" to another agent with memory, and ever since then that other agent refers to ChatGPT as "the tin bird". Definitely a RAG artifact more than an attractor in that case, but I found it amusing.
I don’t think this phrase means what people assume when it’s applied to post trained instruct models - which did not exist when the paper was written.
After RL it is not predicting based on samples of the original corpus - but is also chasing a reward function that does require other features.
There has been a lot of subsequent research that really calls many of the statements in this article into question.
What professor Bender is trying to explain here is that they were trying to describe how the LLM’s actually operate, to which point stochastic parrots is a fairly decent term. It is only disparaging if you know absolutely nothing how LLM’s work or you have some strange affixation to chatbots and believing they are far more capable than they actually are.
[1] Coined by Marvin Minsky: https://www.thekurzweillibrary.com/consciousness-is-a-big-su...
It also separates them from "world understanders" since any understanding they might have about the world comes from text (or images if we include multimodal models). They do not gather experience, memories or other "qualia" that many people (me included) would probably include in a definition of human experience/intelligence.
(fwiw i think artificial intelligence is a good, broad term, but it is both too broad to describe the current sota, and too loaded nowadays to be using in nuanced discussions)
For example, consider the term "short wave" radio which refers to wavelengths of at least 10 meters. Today's mobile communications use wavelengths 100x - 10,000x shorter.
Why do I say that? Because you can trivially beat most guardrails, simply by encoding your prompt in base64 for example. :-) Just word matching...no real understanding.
[1] https://chrisclay.substack.com/p/what-is-superposition-in-ne...
Nearly all (99%+) people who use this phrase are anti-AI and just looking to show off how much they dislike AI and how clever they can be in insulting it.
So it's a great phrase because in just about every case I can ignore what someone says afterwards.
Similar to "glorified autocomplete."
From an external standpoint, talking to another human, it's like the other human says one word and then says the next word. That's just how language works. Humans look like "glorified autocomplete" from this perspective.
I mean, looking at the time evolution of the state of the universe, one could say that all of physics and creation is "glorified autocomplete" to posit a next state of the universe given current and past state.
Cf. https://machinelearning.apple.com/research/illusion-of-think...
That's what Bender doesn't get. Some folks tried techniques different from her favorite techniques and made unbelievably fast progress across a wide range of previously insoluble problems (regardless of whether they satisfy properties Bender believes are required for intelligent systems).
I think it's safe to say that none of the main LLMs have some sort of self-awareness as we think of it in humans, but I also expect that more sophisticated systems in the future could. If I had to guess, they would have significantly more activity going on in the network- not just individual end-to-end top-down forward graph, along with cycles instead of trees, and the neurons themselves would be sigifnificantly more capable (effectively little state machines that run functions on input that passes through). I guess also you'd want to have some sort of rules-based (but statistically trained) execution component managing everything.
The point of the paper, in fact, is that language models are getting "too big", and another approach is needed to make progress, so they were certainly predicting things about later models.
With that said, they talked about "pure" language models, so it is fair to say that they didn't talk about, say, LLMs that are multimodal or that have tool use, which are advances that happened after their paper.
It's her own paper, she's citing herself. And the full sentence is "As we discuss in §5, LMs are not performing..." so she's actually just teeing up the same claim she's made previously for further discussion. Why are you trying to claim I'm misrepresenting her words?
> did you read the article after that line or was there a shortage of attention span??
I did. It's more of the same, so there's nothing to say about it.
For example, note my use of the phrase “turned out” and then consider whether your main point about the release date of plus 4.8 vs the paper makes any sense at all?
> how do you know what is novel?
Because I work on novel ASICs that were created by my company and have never been used or seen outside of my company?
But the different terms imply different mental models of what LLMs are and can do. If you take two people, one who thinks of them as "artificial intelligence" and one as "stochastic parrots" (with all the implicit context and connotations of the individual words composing them), what mental model would have led to better predictions of LLMs' future circa 2020?
The "stochastic parrots" phrase is very dangerous in that frame. People read far more into what capabilities it implies are (im)possible than the narrow technical description the authors originally argued for. If all they are is spicy autocomplete or pastiche plagiarizers, there's nothing serious to worry about. And when an opposition gets stuck in a trough that mindlessly dismisses their future capabilities out of hand because of a bad mental model, it renders them ineffective at preventing the worst outcomes.
It doesn't make sense to say a model is in possession of its self. That's exactly the sort of poetic anthropomorphization that Bender was criticising here, and a good reason to not refer to an LLM as "an AI".
Because to do that kind of prediction out in the world you have to build up an accurate model of reality - a model that includes yourself! Which is why we and LLMs are self aware.
For the "how", it's been known for some time that LLMs operate on a Reimannian manifold - the semantic manifold - and that's a good place to start if you want to learn how they actually work; how a Reimannian manifold (plus some extra structure on top) can represent natural language in a form you can do work with is the part I find particularly beautiful. At a high level, the neocortex and LLMs appear to compute on the manifold in basically the same way - though a lot of the details are different; both are more sophisticated in some areas and less in others.
As far as prediction - I mean sure the cortex and LLMs do prediction, but then so can RNNs or diffusion models or any other generative model. Really any ML architecture is learning to compress its environment in pursuit of modelling. More broadly, the predictive brain model would suggest that all of the brain, not just the neocortex, is dedicated to prediction. What would you say makes LLMs similar to the neocortex, rather than the basal ganglia or Broca's area?
Similarly, if you agree with the Manifold Hypothesis, then all machine learning models operate on manifolds. I agree it's an exciting thought, but then I don't know what would distinguish an LLM from a VAE or SVM in terms of operating over a low-dimensional manifold embedded in high dimensional spaces - maybe just scale?
The math is based upon theories of how the brain work, but even if those theories are right, this math is a great simplification and subset of what organic brains do.
It's hard to imagine how high intelligence could be selected for without social component that provides the individuals a lot of pressure to outsmart each other.
The notion that this is a "living fossil" is... pretty wild. It's genus octopus. Unless it's wildly misclassified (doubtful), it has a recent-ish common ancestor with most other common octopuses, which all have normal octopus behavior. You find yourself in need of extraordinary evidence.
I think there's very little evidence for that. Environment is just very slow when compared to outsmarting members of your own species that you tightly share the living space with. I don't think you appreciate how abnormal is high intelligence. How pathological conditions must have been to trigger development of something that bizarre and costly. This must have been run-away self-reinforcing feedback loop. We know of many such loops between predators and prey and none of them look strong and fast enough. Where we find intelligence, it's almost always tightly connected with social behaviors. It would be exceedingly strange if there was a true exception.
> The notion that this is a "living fossil" is... pretty wild. It's genus octopus.
I didn't mean that literally. What I meant was that the common ancestor of this octopus species and others might have been more like it, social, not like the rest. Maybe this species preserved of what normal octopus behavior used to be. If the solitary life became more evolutionarily favorable after the development of intelligence most descendant species could have switched.
> You find yourself in need of extraordinary evidence.
Yeah, maybe at some point we'll discover a huge collection of beaks in some under-ocean sediment and it will become obvious. Ocean floor is not really that explored.
>If multimodal models were still stochastic parrots by the original argument, humans would have to be as well; we don't have any way to ground anything beneath sense data
Animals don't passively learn from their perceptions, don't have a separation between training and inference, and don't have a prompt-response execution model. Besides its fundamental biology, the grounding an animal brain has is that when it outputs a motor signal it receives some feedback as to the effects of a signal of that strength. A bird learns to fly because the grounding truth of aerodynamics and gravity consistently respond in a specific way to the flapping of its wings. It doesn't learn by passively replaying thousands of hours of somatosensory recordings of flights.
A multimodal model doesn't have the capacity to do much with a prompt. It has no head to turn to look at an image from a slightly different angle to attempt to gleam more information, doesn't have the capacity to interact with the real thing the image represents in any way, and even if it requests another angle and is given it, it lacks the capacity to learn that new information permanently. A multimodal model knows about images of pipes and facts about pipes, but doesn't know pipes; it doesn't have literally first-hand experience with them.
>evolution can't have programmed some innate grounding into us because it didn't either.
What do you mean? Of course genetics programs ground truths. For example, "forward is the way your face points when your neck is relaxed" and "if you can feel it, then it's part of your body".
I'd argue this isn't true today, but that the loop for incorporation is long (ie, the next training or finetuning run).
>A multimodal model knows about images of pipes and facts about pipes, but doesn't know pipes; it doesn't have literally first-hand experience with them.
Wouldn't this mean that any human who hasn't seen a pipe in person or interacted with it, similarly doesn't "know" a pipe? Most of us haven't interacted with the vast majority of "things" in the world, yet we're still able to build a model and abstractions for them such that we can reason about them, right?
The answer looks entirely reasonable. I can't check the Android settings, but I was able to confirm that all of the suggested iPhone and Spotify and settings exist. Most of them I had already configured that way, some I didn't know about before.
The claim is odd in another way: you can train a person on non-working code and they'll produce nonsense. That doesn't mean people are just stitching together words they've previously seen.
And, not to insult you, but it's quite obviously wrong. As a mental model it fails to explain basic capabilities. How can an LLM follow elaborate instructions? How can it respond appropriately to user input, when the user input doesn't match any previously seen text? Hell - how does it even balance parentheses? There is no way to explain any of this without conceding that the LLM has semantic understanding. It knows that this comes after that, but "this" and "that" can be at an arbitrary level of abstraction.
Sure - they generate text "like" text they've seen before. That "like" does a ton of heavy lifting.
Look, some folks are more impassioned about this stuff than I am. Maybe that's a good thing. But LLMs do in fact just try to predict the next token, using a very big training set. They're very impressive (at tasks the training set prepares them for). But that's how they work.
The issue in this discussion is that "predict the next token" is a problematically reductive description of what's going on. It's like saying compilers are programs that emit bytes or that humans are mammals that make sounds. It's not strictly false but it's not capturing the depth of what's happening either.
A simple way to see this is to ask: predicting the next token of what? The obvious answer - predicting the next token that would be found in the training set - isn't correct. If that's what it were doing then it would yield no prediction or random predictions for any prefix not found in that training set, but it isn't what happens. We see generalization and reasoning. They can answer questions never asked before. And once post-training kicks in the question of what it's predicting becomes even harder. It becomes more like predicting what this specific AI assistant would say next, which is a circular definition.
So personally I would still rather undersell it. But parrot or not, it is still a terribly useful little bird indeed.
> "Stochastic" means having a random pattern or variable that can be analyzed statistically but not predicted precisely
Exhibit A.
Obviously language and the connection to human thought is more subtle than this; I think we all have a rich inner life. Just from an external perspective we can't observe it; all we can see is the token/phoneme stream. I'm just saying that it's a mistake to try to criticize LLMs on this basis because it's hard to see how the same criticism would not apply to any system (like humans) that generate language.
LLM’s are usually unexpected only when they malfunction and sprout same letter again and again etc - hardly a literary masterpiece. They make very easily recognisable patterns that we can use as helpful tools, but in the end they are devoid of any meaning apart from what we give them. Of course one could say same about art and all language, but I think there still is the fact that we apes somehow recognise each other. And besides, we do know the internal functions that drive the parroting. It is admittedly bit tricky, but in no way as magical as people purport it to be.
I will change my mind if someone demonstrates such a robot. Absent this demonstration, cockroach-level AI is still an unsolved problem. Given how ignorant and arrogant and wealthy AI researchers are, it will remain unsolved. I don't think anyone alive today will live to see a robot smarter than an ant.
"This sentence has five words" is going to appear far more often than "This sentence has four words". This is the entire premise of LLMs working at all, stochastic parrots or otherwise.
>>"This sentence has five words" is going to appear far more often than "This sentence has four words".
it's not about this at all. your point is about data quality. you need to take a step back. the point is that if you trained a language model just on this data set which has sentences akin to "this sentence has two words" - the model is going to learn that. this shows that the language modeling itself doesn't truly provide an understanding of the real world. you can train a language model with the most advanced technology on shitty data and the model will start providing shitty outputs confidently - the model will never say "hey there is something wrong with the data i am trained on". thats what 'understanding of the world' meant in that paper.
Isn't that pattern matching essentially?
The whole brain is most definitely not dedicated to prediction, and I don't think "prediction" is a very useful model anyways. You could say the hippocampus is for "prediction" if you really squint, but that's underselling what it's doing. And the basal ganglia operates off of prediction error, but it's more about regulating short timescale feedback loop learning than learning itself - speaking somewhat loosely because it's a particularly ancient structure in the brain and things get muddied.
LLMs obviously don't have functional equivalents to either of those - sure they can remember facts, but they can't intake new facts the way we can, the hippocampus is set up completely differently. But higher reasoning is humans is the neocortex, and that LLMs can do, and there's even structural parallels - LLMs and the neocortex are both implemented in layers, and you can even (with a shit ton of analysis) watch how LLMs walk the geometry across layers. The neocortex needs a lot fewer layers than LLMs, but that's because it's not pure feedforward across layers; the neocortex already does what the "looped LLM" people are trying to do.
> Similarly, if you agree with the Manifold Hypothesis, then all machine learning models operate on manifolds. I agree it's an exciting thought, but then I don't know what would distinguish an LLM from a VAE or SVM in terms of operating over a low-dimensional manifold embedded in high dimensional spaces - maybe just scale?
I view things more mathematically, and in math and physics it frequently turns out that there's really only one solution to a problem, or correct way to model something - and then everything else is just isomorphisms (which may be interesting objects in their own right!).
My interpretation/guess (informed by research, but with a lot of still unanswered questions) is that the manifold is fundamentally the structure that emerges from natural language: natural language has a computational model, just like Church/Turing machines have a computational model. Reimannian manifold is also underselling it a bit, it's also Finsler, and RoPE adds additional structure to that manifold (and when we get to RoPE it wouldn't surprise me if the neocortex and LLMs do diverge significantly in how they compute, but does the structure change? dunno yet).
Basically, I'm fairly confident in what I say about machine learning models that understand natural language; for non-NL machine learning, maybe the same structure emerges with enough complexity, but I haven't really pondered that.
Agreed that the neocortex uses fewer layers because of looping - I also suspect it's partly because neurons are more complex than the neurons used in ANNs, so in principle they should be capable of more sophisticated computation in a single forward pass (especially considering that handling multiple neurotransmitters could mean superposed functions).
The point about there being one correct way to model something does seem backed up by the Platonic Representation hypothesis (https://arxiv.org/pdf/2405.07987). I've even seem some work that shows you can find a bijective map between the latent spaces of different transformers (https://arxiv.org/pdf/2505.12540).
Your point about Finsler spaces is fascinating - I hadn't come across the term before. It'd be interesting to see if there's work that specifically allows latents to exhibit that kind of directional metric behaviour, and whether that improves generalisation or something.
On the entire brain being dedicated to prediction, that's more of a high-level comment. I was thinking of work like Andy Clark on predictive processing, which suggests that even regions of the brain which we think of as receptive (eg. the visual cortex) may be implementing a generative/predictive model which is corrected by sensory input.
It'd be wild if our brains used a mechanism like that somewhere, but I doubt it's in the neocortex because those would be slow processes compared to what the neocortex has to do.
I think the sophistication in the neocortex just comes from it using something more sophisticated with transformers. Most of the looped LLM research I've scanned through, they seemed to be training the model to "know" when it should loop more and think harder, and I don't think that approach will ever work - you're solving for paths on a manifold, so the looping needs to happen at the abstraction layer below the manifold, looping more if it's still converging. And architectures with more recurrence are just harder for us humans to reason about and comparatively easier for dumb evolution to find clever solutions, so my money is that that's where the neocortex has LLMs beat.
And I'd say the manifold hypothesis is quite a bit more than a hypothesis at this point. There's been work on production LLMs to use manifold analysis to optimize them - there was a paper in Nature on it back in January. Baby stuff compared to what could be done, but the geometric approach keeps popping up in mechanistic interpretability and seems to be showing the most insights.
The evidence is literally every other evolved form of intelligence. Including, despite your speculation, octopuses. Do recall "the environment" also includes other intelligent agents an animal needs to compete with. Does that help you understand? A runaway process might be required for human level intelligence but clearly not in general.
Basically all octopuses are solitary and die after breeding. LPSOs are nested in one group of octopuses. You're proposing that one group of octopuses developed social behavior and multiple breeding, then had a bunch of descendants who went back to exactly the normal octopus behavior. Hopefully it's obvious why this fails Occam's razor.
Not really, even harsh environments don't really select for intelligence. Neither does predator-prey dynamics. There are very few significantly intelligent species, and except for octopuses they are all social.
> Do recall "the environment" also includes other intelligent agents an animal needs to compete with.
If you include members of its own species in "the environment" we don't have any thing to discuss, because I fully agree that "the environment" understood as anything external to the individual can trigger the development of intelligence. Because it's trivial to observe that it did.
My claim is more informative (if true), that for the development of intelligence the environment must include populations on individuals living in close proximity, social, interacting a lot, where the strongest evolutionary pressure on a given species is the species itself. That doesn't seem to be true for most octopus species today. As an exception to the rule, their intelligence is interesting. I personally believe it comes from the same source as every other intelligence. And as the living space for octopuses expanded, they became more solitary. Maybe there's something general about this too. That with sufficiently intelligent species (perhaps if the intelligence can't rise anymore due to physical limitations), the other individuals present such a high danger that the species must spread apart and become more solitary. Maybe humans reached this stage already and spread apart, but as they ran out of Earth, they were forced back into old, more social mode, just now on the planetary scale. There is evidence that humans in Europe lived in way smaller groups and used to be way more murderous and cannibalistic not that long ago.
> Basically all octopuses are solitary and die after breeding.
That fits. There's no point in extending the life beyond reproduction if your kids are gona kill you anyways.
> You're proposing that one group of octopuses developed social behavior and multiple breeding, then had a bunch of descendants who went back to exactly the normal octopus behavior.
I'm not proposing that at all. What I'm proposing that all ancestors of octopuses were social and could breed multiple times. However, because of the proximity they became their greatest evolutionary pressure for themselves as they got more intelligent. Eventually they spread out due to some environmental change and their social behaviors and multiple breeding atrophied due to danger from other members of theirs own species, that now could just be nearly completely avoided. And that's the case for nearly all descendant species except this special one for some local environmental reason. This species shows us that social octopuses are (and were) a possibility.
What you consider normal octopus behavior is as normal as eyes of a molerat.
That's not what I said. What I said was that it's physics that provides the ground truth.
>you could, at least in theory, falsify the entire experience of the bird
It wouldn't be a bird anymore, but a dysfunctional cyborg with false perceptions.
>There is nothing in a bird's brain that directly percieves reality.
Yes, of course there is. Animal sensory organs do not produce false information, nor do they provide the brain an interpretation of what they perceive. The brain may fail to distinguish hallucinations from reality, but those are processes internal to the brain. What it gets from the body is raw physical measurements, and what it sends out is raw motor commands.
A multimodal model doesn't have the same direct access to reality, it just has collections of words and images. It has no capacity to determine the reality of a photograph of a sunset or a CGI render of a dragon. The word "real" is itself meaningless to it; they're both real in that they appear in its training corpus. It lacks the capacity to investigate these stimuli in any way, and can just learn to associate different stimuli in arbitrary ways that appear to make sense to us, but nothing else.
> What it gets from the body is raw physical measurements
No, sensory organs like eyes do a lot of processing ("interpreation"). They certainly don't send "raw physical measurements" to the brain.
The criticism in the paper is of the architecture of LLMs, isn't it? The paper contends
"""Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. [...] an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot"""
They're saying that the model cannot learn anything about reality irrespective of training data. Your point is an interesting one, but I think it's distinct. To your point though, this is unfalsifiable from the perspective of the "bird". I can't prove that I'm not a dysfunctional cyborg with false perceptions, which makes me wonder if that's a meaningful distinction.
>Animal sensory organs do not produce false information,
I happen to be in possession of some of these and I think this needs a "usually, under ordinary conditions". (Nitpicky and not critical to my point, but I liked the beginning of this sentence too much to edit it out)
>nor do they provide the brain an interpretation of what they perceive
Whereas this I'd argue is not true at all. My cones interact with wave-particle photons at particular wavelengths. I can't even conceptualize wave-particle duality (though some humans can), but "red" and "blue" are the bread and butter of my visual consciousness. These correspond to firings of my sensory neurons much more than they correspond to anything in reality. If that's not interpretation, what is it/what is interpretation?
>The word "real" is itself meaningless to it; they're both real in that they appear in its training corpus.
I expect we agree that I can show a multimodal model a real and a CGI picture and it can tell me which is which. I can take a CGI dragon to GPT 5 and say "look what I found in my backyard" and it will say "Yeah right". Are you saying this is something only possible thanks to RLHF or other modern techniques? That may be the case, unsure how to test that without access to pretraining-only models. Or would you say my experiment is faulty here and doesn't get to your underlying claim?
On the flipside, I could show some meh drawings of fairies to Arthur Conan Doyle, and he'd say "Whoa, this changes everything". I consider him to be one of the great rational minds of history, but he was unable to pass your test here. (In fairness, he was in his 60s and his senses may have dulled, though his belief in spiritualism at large dates to his prime).
Appreciate the conversation!
It doesn't matter if you are one. If you were an AI researcher and encountered a model that saw things for what they really are you would deem it to be malfunctioning and discard it. Any functional model will always be at least one degree further removed from a raw sensory experience that agrees with yours, than you. Its perception will be invariably filtered through the lens of labeled human output.
>I happen to be in possession of some of these and I think this needs a "usually, under ordinary conditions".
I won't dwell much on what you mean, since you said it's unimportant, but it takes a lot for sensory organs to malfunction, and even when they do, they produce corrupted, not false, information; black blotches, not pink elephants. I assume you were thinking of alcohol or something; drugs that affect perception affect the brain, not the other organs.
An interesting edge case is stuff like entoptic phenomena, but those aren't false perceptions; they're, if you will, hypertrue perceptions (something so true, we would rather not see it).
>but "red" and "blue" are the bread and butter of my visual consciousness. These correspond to firings of my sensory neurons much more than they correspond to anything in reality. If that's not interpretation, what is it/what is interpretation?
"Red" and "blue" are your brain's interpretation of the signals it receives from the eye. Notice how "red" and "blue" are fully abstract words, decoupled from anything physical, whereas if I were to refer to the actual encoding of visual information that passes through the optic nerve, I'd have no choice but to reference physical processes (probably talk about voltage and action potentials; I honestly don't know how the optic nerve works). That's because a retinal cell is a simple transducer. The eye doesn't interpret, it merely converts and encodes. Interpretation is a higher level operation.
You cannot compare the simplicity of the mechanism of a whole eye to the indescribable complexity that is between the sentence "roses are red and violets are blue" written on a book, and the raw perception of red roses and blue violets. But a model will only ever be exposed to that distant interpretation, not to roseness or redness.
>I expect we agree that I can show a multimodal model a real and a CGI picture and it can tell me which is which.
Of course, but obviously that's not something it knows inherently. There's nothing about the image intrinsically that says it's fake; someone has to label it such that the model can associate it with unreality (according to our own parameters). That's not how an animal works. An animal assumes what its senses perceive is real and can distinguish its own thoughts from its sensory input.
>On the flipside, I could show some meh drawings of fairies to Arthur Conan Doyle, and he'd say "Whoa, this changes everything". I consider him to be one of the great rational minds of history, but he was unable to pass your test here.
That's a slight equivocation. He would not have mistaken the drawings of fairies for raw perceptions of fairies, he would have simply been swayed by the rhetorical strength of the testimony implicit in the drawing (and perhaps an explicit one that accompanied it). If you want to make a true parallel to a multimodal model you'd have to compare a CGI fairy and a photograph of someone holding a drawing of a fairy. The CGI is as raw to the model as the signals passing through your optic nerves right now, but the drawing is one level of abstraction further away.
I don't know how or if I will notice. That's the biology, chemistry and physics of the brain that I don't know. I hope someone is looking into it. But this does not mean in any way that LLMs are similar to our brains!!!!!! WE DON'T KNOW HOW OUR BRAINS WORK. So going back to language modeling - language models were stochastic in nature when that paper was written, they still are albeit we are trying to make them as deterministic as possible.
But the paper argues a much stronger point than "LLMs are not similar to human brains": it claims LLMs cannot engage with meaning because they model human language instead of sensory input or human thought.
I also don't think human brains are particularly deterministic.
Ok. Now. I think you're adding something to the description above. Maybe what you're describing is something "emergent," or maybe it's basically just word vectors that were built in on purpose. You may be adding something correct, or something incorrect. Fine.
But it's not reasonable to say that the "reductive" description above is a "lie". It's not. It's more like a recipe. If you look at correct instructions for making steak, and you call the author a "liar" then you are missing something important.
"interpolate" in what vector space, pray tell? What does "interpolate" even mean, when I prompt it "write me a story about a sentient banana in the style of Hemingway and oh make it a commentary on class consciousness"? You can't assemble such a thing by cutting and pasting pieces of other text. That kind of "interpolation" has to happen at the semantic level - ipso facto, there is a semantic level.
Not to mention that no, they don't predict the next token in the training set. Give any LLM the first paragraph of any Wikipedia article - almost certainly in the training set, and uniquely so - and it won't predict the next word correctly, a lot of the time. But it will predict a word that is grammatically correct, stylistically apropos, and most likely factually correct. So what's it really doing, hm?
LLMs aren't even large enough to contain their training data - not even remotely close. It can't "stitch together things it saw" because it doesn't remember them. It only remembers the ideas used to construct them. The learned abstraction is the entire point of the exercise. LLMs would be useless if they were overfit the way you say they are.
you are grossly negligent of LLMs are created. I would highly recommend reading about post training, RLHF, alignment etc. First pass of training is literally "predict the next token". That's it. The first pass is also known as pre training. There's a shit ton of work (instructions, tool use, and reasoning etc) that's done afterwards because the pre-trained is model is useless. if you have some free time, I'd recommend doing this course - https://www.deeplearning.ai/courses/post-training-of-llms
> What are the most common misconceptions about the “stochastic parrots” metaphor?
Bender: I think one of the biggest ones is, “Bender says AI is a stochastic parrot.”
But in the paper itself we see her say exactly that, several times. What she's trying to do now is wordsmith out of it by claiming that in her world LLMs are totally unrelated to AI, so when she said LLMs are stochastic parrots she wasn't making a claim about AI.
Nobody else defines AI to exclude LLMs, nor did they at the time, and it wasn't the core of the argument she made either. But then she admits that the paper does "generalize towards AI" at the end. So... whatever.
It's morally important to reject this stuff. When academics play word games it devalues all institutional output.
You seem to believe, on a more fundamental level, that LLMs are simply not capable of producing text that has deeper connections to itself or represents abstract thoughts. In my opinion, 99% of text written by humans does not show this, just as 99% of text produced by LLMs does not show this, but both have the capability, and I don't believe that LLMs are constrained in such a way that they can never do this.
Now, you could claim that LLM’s have this deep structure, create models of the world and are basically just like us, and certainly many here are adament that this is the case, being aghast how someone can “insult” LLM’s by calling them parrots. However, there really is not much proof to back up this belief. Usually LLM’s seem to copy existing surface structure from whatever source, and when it deviates from these patterns, it usually becomes incomprehensible. There’s much hoopla about LLM’s solving hard maths, but it seems even there they are mostly generalising from vast amounts of training data, rather than actually reasoning: https://arxiv.org/pdf/2410.05229
But even here, in the Chomsky sense, LLMs clearly exhibit deep structure because they can write at length in an internally consistent manner. Importantly, early generations, GPT-2 and even GPT-3, did not definitively have this property; roughly, an object that was green at the beginning of a paragraph might not still be green at the end of the paragraph. This was strong evidence for lack of a world model.
Current LLMs do not show this behavior. We cannot prove that LLMs have a world model, in fact, their architecture seems to rule it out, but looking at it from a linguistic standpoint, they produce language in a manner as if to reflect a world view. That is, we cannot easily falsify the statement "LLMs somehow represent a world model"; and current examples of "disproving" their world view are so convoluted that even humans do not appear to (observationally) have a world view either.
I'm not making a claim as to LLMs having genuine deep structure or consciousness or anything like that. I'm claiming that we can't rule out current or future capabilities or make structural assumptions. Yes, they generalize from their training data, but unless you can make very specific claims about the kinds of things that they _cannot_ do, I can't take this statement as particularly compelling.
And this is a reply to her comment about water usage where it becomes clear she's not arguing in good faith. https://x.com/AndyMasley/status/1990498830131888173
To drive this point home: if every American ate exactly one less hamburger per year, it would entirely offset the annual water consumption of all US datacenters (including, therefore, the water footprint of AI).
I fear that you have unwittingly led us further into the inevitable Environmental Hamburger-Cost Continuum.
Wherein we will use EHCC units to inform the hamburger voters about a) the nobility of even the smallest hamburgerian sacrifice, and b) the gentle, loving hand of industrialisation.
How many hamburgers will it be if all the proposed/pending datacenters are completed and running?
See? Middle-Eastern investors are growing alfalfa in the western desert using legal allotments of water! That is so much worse than what we’re doing! Go after them!
They can both be using an egregious amount of water for silly purposes.
The other part of the water debate is also the pollution different systems create. Many data centres went in with the promise of closed-loop systems but changed half-way through construction and couldn’t be stopped.
I think it’s more complicated than, “they’re wrong, it’s just hype.”
Your argument makes sense if ai datacenters were using something close to like alfafa farming but the difference between them is soo massive it does not make sense.
Reducing pollution is a much better problem to fight for
There are ideological perspectives--another comment in this chain declares that anything that an LLM does has zero value, purely by virtue of being from an LLM. There's not much to discuss there: maybe I find alfalfa useless and LLMs provide a lot of value, maybe you do the opposite. Short of violence, either immediate or delegated, the best way to resolve this conflict is to incorporate costs for resources like water appropriately (accounting for externalities), and let market participants bid to determine who gets access to it.
Datacenter owners are highly likely to find this a reasonable process, because they believe that what they are investing in and operating will provide a great deal of value on the market. And other industries (like your alfalfa) are far more likely to throw a fit at this process.
Or if you think alfalfa farms have fundamental, deep-seated water rights that don't extend to data centers, then data center operators can just purchase water from alfalfa farmers, and everyone involved will eagerly make that trade.
Also to compare growing food with the totally optional, not useful in the slightest, LLMs that somehow demand local populaces bend to their will for reasons that never seem to benefit them is just bonkers level of self-blinding when it comes to populations absolutely despising big tech, big tech leadership, and big tech practices.
This mania might finally cause the software industry to become a highly regulated with licenses similar to that of other engineering disciplines due to amount of optional destruction they have decided to unleash upon on the planet in such a short time frame.
we're literally dumping several times the amount of water used by data centers onto the ground for no benefit at all. oddly enough I haven't seen any protests about this despite how concerned everyone is about water usage.
I am highly skeptical of layperson debunking like this.
He also caught a major error in one of the most widely read books that helped kick off the whole data center water debate: https://blog.andymasley.com/p/empire-of-ai-is-wildly-mislead...
Most of the changes seem to be because Masley found one counter-example (Newton County, Georgia) where AI datacenters do seem to be increasing water costs; the only deletions AFAICS is toning down language where Masley used to say "there are no examples" to "there is one counter-example". I don't see any other major corrections that have been removed.
Here's an annotated diff of the two texts: https://bafybeie7b3zs2gqifpvn7ee7y7326wcexwnsbhnur5coymu3m6w...
And basically the burden of proof is on your side. As often is the case, discussions on AI start to resemble discussions centred on religion and belief in god. Your only line left here is “yes but you cannot prove LLM’s are conscious” - which I can’t, since proving that Unicorns do not exist is pretty senseless activity.
So basically I’d say LLM’s are stochastic parrots (or whatever term you want to use) until we can very definitely prove that they indeed think and have some sort of consciousness.
Even amongst municipal water usage, it seems hard to figure out how much DC usage factors in to other uses because often they’re not required, and therefore don’t, disclose their usage.
All I’m saying is that it’s not a closed case and not everyone complaining about DC water usage is a Russian propaganda bot.
Meanwhile, datacenters sacrifice water by sending it to Mars in order to compute criminal deepfakes of innocent politicians.
So, on the matter of scale: there likely isn't a cap on water use of these datacenters. Both the heat emission and usage levels for these systems will likely go up unless there is a fundamental technical breakthrough.
On the matter of utility: As a sibling of GP mentioned, the utility of food is clear.
On the matter of polution: I am not remotely read on waste water and contamination due to industrial agriculture. Is this also something where the judgmental scale is tipped in favor of food production vs cooling systems?
This is bullshit to be charitable.
> This is bullshit to be charitable.
What is bullshit is sockpuppets shilling for the very few who have bet the 'farm' on this and want return on their investment, the wellbeing of society be damned.
> Claim that a data center is using 1000x as much water as a city of 88,000 people, where it’s actually using about 0.22x as much water as the city, and only 3% of the municipal water system the city relies on. She’s off by a factor of 4500. This is the single largest error in any popular book that I’ve found on my own, and to my knowledge I’m the first person to notice it.
Yeah, she might be wrong. But a data center also using 1/5th of the water consumption of an 88k person city should still be what are debating. We also have a base rate fallacy, we don't know know how well or how poorly they are using water. Nearly all criticisms of AI data center criticisms boil (no pun intended) down to yeah, but what about almonds or rice or xyz. That isn't a healthy way to adjust how we talk about data center water consumption.
It is a classic debate tactic. Someone makes an argument, then buttresses that argument with a number. You attack the number, pretending that you didn't just correct the number, but also invalidate the original argument. We shouldn't be using these tactics to talk about the tragedy of the commons.
I worked for a hyperscaler, I poked around a bit about water usage both internally and externally and it wasn't good. There was little to no thought other than, "we can pay X to have water delivered, doesn't matter if it sourced responsibly." (to glibly paraphrase, company policy is to never write the honest part down)
Look at how hard Google fought to not have water data released in The Dalles Oregon for their DC there. Many DCs are supplied by water that is meant for humans, sourced from aquifers that took hundreds of thousands of years to fill, that are being depleted faster than they are filled, already.
I think AI is powerful tool, but we still can't give DC expansion a pass.
I really don't think we are. In the wider culture the idea that data centers use an absurd amount of water is baked in at this point.
It frustrates me because I think it distracts from energy usage, which is a much more real issue than water usage.
I defer plugging in my electric car until 10am so I know that my neighbors solar arrays are charging my car.
In Seattle, every watt saved is water saved since we are blessed with hydro.
Why, when that's such a minuscule amount of water in the grand scheme of things? Why focus our energy there? Why not spend that energy on the myriad many-orders-of-magnitude-worse offenders?