LLMs are not suitable for brainstorming(piaoyang0.wordpress.com) |
LLMs are not suitable for brainstorming(piaoyang0.wordpress.com) |
The problem with this argument is that people do the same thing, we’re not that great at brainstorming either. When we brainstorm in groups, we’re just bringing multiple points of view together. The more data LLMs are trained on, the more viewpoints it might be able to bring that you haven’t considered.
That said, LLMs and all NNs so far are built to interpolate, and they are bad and have unbounded error when extrapolating outside their training examples. That is a good reason to not expect today’s AI to come up with new ideas.
You're right, most people do that, but that's because they haven't trained themselves to be inventors. This is a skill that definitely needs to be developed more, but finding truly creative (and sometimes backwards-seeming) solutions can sometimes mean thinking about problems in a completely different way.
However, all LLM's do this. That's one of the points of the article.
LLM’s degree of adherence to training patterns can be directly tuned, per request (up to whatever the maximum it is capable of.)
Can this be leveraged to build a framework around an LLM for brainstorming usefully? I’m not sure. But “Nope, because LLMs just directly adhere to existing patterns” is definitely not a useful answer to that question. It is the kind of answer someone who just parrots common patterns with surface-level knowledge of the subject area would come up with. (Huh, did a default-settings LLM write this article?)
I don't think handwaving about exclusively outputting training data and random punditry via article really holds up here.
It's wrong for so many reasons, from the training data having more perspectives, to the "Fat Tony" test, empirically, we can go make it output novel combinations right now, to the CS test, we can have it emulate arbitrary Turing machines.
As a part-time generative artist for several decades, a user of Monte-Carlo methods on a daily basis, and author of some papers on the topic, I personally believe that randomness is not a good answer to anything creative. Randomness is boring, and average. Randomness only helps you when you have a high quality Markov model constraining what your RNG chooses between, and that’s more or less all LLMs actually are. Adding more randomness to creative works in general makes creative works muddy and lowers quality, it needs to be guided. Randomness is a useful tool, but is widely misunderstood IMO and not very effective at brainstorming style exploration; brainstorming is about solving problems in interesting ways, i.e., there is reasoning behind it, not slightly more random word salad.
how do you know?
When I'm brainstorming math or cryptography or philosophy with an LLM it is only because of the connections and leaps that I introduce that leads to novel and amusing discoveries.
You have to develop successful patterns. Coaxing and condensing. Drilling into numbered lists. Asking for tables with progressively more columns and altered sorting. You have to know the spatial reasoning limits of the model. Asking for enumeration and disabling of social niceties. The number of skills needed is probably beyond most people's imaginations.
I would say, try reading a conspiracy theory forum - or simply wait for one to get posted on HN, and then revisit this particular conclusion.
I don't think this is a unique failure of LLMs, except insofar as LLMs have less of a physical grounding in reality then humans do: but we have people out there who have convinced themselves they can project forcefields, and take that all the way through to getting punched in the face by a martial artist when it turns out that, no, they can't.
I have often inadvertently found myself brainstorming just by explaining something challenging I am working on to a colleague with no expertise in my area. Just explaining forces new perspectives and new ideas.
Similarly, LLMs are captive audiences for throwing out ideas and playing with them. The interaction creates something more than just talking to myself or a whiteboard.
Not all brainstorming partners need to be first order contributors to have a very positive impact. Different kinds of interactions can spark entirely different kinds of ideas.
If you've never talked with a highly opinionated LLM, I understand your scepticism, but there is certainly value if you use the right models creatively.
Or even like "huh? explain"
Looking forward to ramping up the honesty parameter. Hoping OpenAI's new voice model trivializes this so I don't need to prompt engineer.
Chargpt: “that’s a great idea for a business! Here are some suggestions…”
Not a real exchange, but god, it does feel that way sometimes. One reason you can’t trust the damn thing is it’s too positive.
I want to acknowledge that the original title is inaccurate, as many of you have pointed out. It should be "(current) LLMs don't brainstorm novel things really well" rather than just brainstorming.
It's not intended to be a clickbait though - I was kind of mixing two definitions of brainstorming unintentionally. When we refer to the group activity that aims to collect all angles from participants (and common wisdoms), LLMs are really good, and it's something I do on a regular basis. However when it comes to the hope of reaching novel ideas that don't exist before (which some of us will consider what distinguishes brainstorming from group discussion or research study), I would say today's LLMs don't do well. I've seen such issue in business and arts domains, and also someone here mentioned similar experience in video game design.
I would argue that (so far) for any idea LLMs tell us, there exists at least one instance of a similar pattern in the training data (either exact or in a high level). If this is what you need, then great. But some problems require more than that. And I would argue that a lot of important innovations in history didn't follow this pattern. I'm aware of reports on LLMs helping research (e.g. the works shared by Terry Tao), but I don't think they contradict the point here. Will be super happy to be proven wrong though!
You submitted the post?!
It got “picked up” so quickly because it is wrong. I think what your post does show is how effectively you can farm HN for supporting arguments.
You provide zero evidence for your claims, nor do you give us any transcripts of your attempts. Every post that matches the structure of yours is usually a summary of how the author has low skill in using LLMs.
Though I am absolutely delighted in the comments and the nih paper referenced was a joy to read.
I'll also add that there's a big difference between "give me 10 startup ideas" and having a conversation with it to explore different startup ideas (for example). Through conversation you can build off what each other say and explore the space.
Was thinking if you can apply TRIZ method to invent Transformer before 2017 - hard to imagine it working on me
LLMs are going to give you the consensus reality (average opinion? average facts?) but you can easily steer it into offbeat, controversial and esoteric areas of it's training with the right prompts
ChatGPT more or less it converged on simply taking the current subject, adding a generic gameplay element to it and outputting it. It took half the ideas being thrown at it and included a rhythm mechanic... That's not that interesting, and I just couldn't get it to really think outside the box.
I put a clip up supporting why I think LLMs are good for brainstorming from Marc Andreeson about this topic in case your interested
"Marc Andreessen says with the right prompting, you can unlock the latent super genius in AI models"
https://www.youtube.com/watch?v=N2yN4IG8UYA
And here's my comment on another site to a pro novelist who suddenly discovered Claude wrote like a genius who posits that Anthropic deliberatly cripples Claude's writing ability because they don't want to scare writers suddenly, they want to ease them into it
I disagree. Your supernatural scenes triggered words an analysis from higher quality writers and commenters in the training data. If you were writing about bass fishing it would likely not impress you with it's writing. You can try something like this as an experiment. In other words it's good at some writing and bad at others, depending on the training data. I'd love to be proven wrong, like a gripping story about bass fishing might be interesting.
LLM might be mostly bad at it, but that doesn't make them unsuitable. They don't have a career to worry about. They don't care what peers or the boss thinks. Etc.
They can spit out ideas - LOTS of them - that real humans then use as a starting point to carry on with.
For more details see "Co-intelligence" by Ethan Mollick.
1. Ask agent to write a few paragraphs.
2. Notice that it's terrible and it would be easier to rewrite it myself than fix it.
3. Actually be motivated to write."LLMs are no good for [use case X]" often means "I aren't very good at using LLMs for [use case X]".
With many powerful tools - violins, say, or carpentry tools - we know that it takes a long time and a lot of learning to achieve competent performance, much less virtuoso performance. Someone who spent ten hours learning the violin and concluded "Violins sound terrible" wouldn't have diagnosed a problem with violins, but with their own mastery. I certainly think current LLMs have some big intrinsic weaknesses, but also that what they are is quite subtle.
I found that og gpt4 is better for brainstorming than gpt4 turbo, which in tern is better than gtp4o.
If you're just using the web based portal you don't get much of a choice which model to use or it's temperature.
No, they won’t come up with a totally new set of concepts and axioms that are totally unconnected to any human set… and you don’t want them to. Because those are the actual hallucinations. New ideas need to be novel, but they still need to fit into the rest of the set.
It’s remarkable just how much people race to be pessimistic about LLMs, even when the pessimistic assertions are obviously and factually incorrect. Like yes we know LLMs have “hallucinations”, which is a loaded term to begin with, because the pessimists literally never shut up about it, even though it’s a fairly trivial overall problem given the immense overall utility.
Some people are clearly so very very emotionally invested in there being nothing productive or useful about AI if it’s not AGI or if it needs the slightest amount of correction or oversight.
How often is that a successful approach at all? I can’t think of a good case study off the top of my head.
The SPC also have a blog that proposes brainstorming from -1 to 1: https://blog.southparkcommons.com/how-to-go-from-minus-1-to-...
Great title, they baited and switched.
95% of the time I don't need effective brainstorming, I need a bunch of ideas, let me pick the best, and move on.
If its a real engineering problem, then I need
>truly effective brainstorming.
They're asking the wrong type of work from it. If you need some boilerplate or a transformation, it's going to give you a fantastic template to work with. If you need it, on the other hand, to engineer out a highly-specific and nuanced solution with an esoteric codebase to a complex problem, maybe not so much. The former is wide, the latter is narrow. It's going to take maybe a bit more breakdown of the scope into proper subproblems before you'll get a good answer; and that's something you can do yourself, or have an agent perform across multiple queries maybe (though I'll admit, more work needs to be done for the whole multi-agent workflows to be truly useful).
Their comment was that ChatGPT can't write final reports but it's extremely useful for seeding brainstorming sessions with real human engineers.
It can generate a list of health, safety, and environmental concerns related to the project that can be distributed and act as a "starter dot point list" that engineers can agree to, expand on, reject as being just silly, and add to as tengential thoughts arise.
I'd argue that makes it useful in the context of real engineering brainstorming .. and that it's being used that way on multi million dollar projects in a billion dollar domain.
When brainstorming this wide knowledge is what you want. The trick is (as always) better prompting to push it hard - so things like "consider parallels in similar situations in other fields" are useful.
Although I might be biased, I love using AI for brainstorming and started building a tool for it https://youtu.be/t5gfETbUzy0
What is creativity if not what people call hallucination? Horribly named phenomenon, btw.
If a coworker confidently gave me code that is not only wrong, but relies on things that don't exist, I wouldn't call it creative. If they tried to claim that no I'm wrong, it is real, I might even say they hallucinated it.
I mean, what is the difference between creativity and hallucination (honest question)?
---
Maybe this means AI would be better at the opposite - after you brainstorm creative out there ideas, AI can tell you if they have been tried in the past and what the consensus view is on why they failed, allowing you to adjust course.
Well, creative problem solving involves getting new ideas and perspectives by creating novel associations between things we know. Hallucinations are creating things we "know" because they sound like they could be right and basing "ideas" on them. In short, hallucinations are bullshit. Totally open creative problem solving isn't always perfect, but even the ideas that don't really work can reveal something about the concepts involved. And bullshit often takes creativity to create, but that doesn't make it useful as actual creative output. It's not like the second you move beyond empirically provable statements, everything has equal merit. If you're trying to creatively work to a useful end, whether or not the knowledge you're working with is fabricated is pretty consequential. Doing otherwise would be like trying to optimize your code based on utterly fake but plausible algorithms-- sounding right isn't right enough to be useful. IMO, being able to create lies that pass the smell test is LLMs' most dangerous proclivity, especially when they're presented as expertise-in-a-box.
I'm not saying llms are good at creativity (i certainly don't think they are) but i kind of feel like its a difference in quality not kind.
Like if you asked me to describe what it means for a human to be creative, i would probably write something quite similar to what you wrote above.
With humans creatively solving problems, what they say isn't 'output'-- it's externalizing a reasoning process. Being off-the-mark about something is a beneficial part of defining the bounds of an unknown solution because it's based on reality. It's a tangible idea that can be reasoned about and modified-- not a collection of output. It's not a point on a scatter plot, it's a waypoint towards a useful end.
Continually asking the LLM to create exhaustive lists of responses to new ideas (big or small) as they first appear, remaining apropos to the running context, is wonderful.
Only one unexpected connection out of 20 or 50 is fantastic, when you can iterate quickly. Brainstorming is the tradeoff of accepting lots of "failure" ideas, in order to generate more serendipitous and creative ideas.
As I use it, the LLM accelerates, augments, and documents my brainstorming session.
--
For some reason, I find it easy to push LLMs to be helpful in ways other people often say "LLMs can't do that".
It is just a matter of setting up the right context, being the right guide, and iterating. They are often faster, aware of more concepts, and more versatile (in terms of fulfilling roles), than any single human.
Their biggest weakness is having a short context. With some work, multiple sessions can mitigate that. Like getting help on a problem from a series of people, where you have to communicate the progression of context to each.
The second weakness is their default to conventional responses. But that is easily overcome by iterative/patient pushing for more creativity responses. Within a session, you can see the model transitioning to more creative as originality goes up, and it gets increasingly "emotionally excited".
They start having "fun", and get more adventurous, as epiphanies emerging from either of you go up. This is not just a funny artifact of how we behave, but also a great barometer for achieving LLM "creative mode".
This causes the LLM to orbit around some point in the latent space, but doesn't cause it to explore. You have to tell it to pick a direction and there is no better direction than a question.
First they have a very small sample size of around n = 150. Second, they don't represent average husband, they use people working on mechanical Turk for $9 an hour.
Finally, they don't have any standard measures of divergent thought - they just use 4 different types of question answering problems, and those are supposed to measure creativity. Whether they do or don't would require us to evaluate the questions.
It's important to be clear that this is saying that LLMs might be more creative than low paid workers on mechanical Turk, given no training time. That's not a very good representation of the average human, or the fact that humans need training.
So for example, an average doctor or scientist might still be more creative than LLMs.
Tldr; this is not a good paper to throw around if you want to convince people. It's truly awful in its design, and the way it's presented is not always honest.
This is a myth. Just because you don't really see blinding insight every day doesn't mean it doesn't exist (and it's more spectacular when you do see it!)
Another reason that this seems to happen is that innovation and invention tend to happen in more esoteric and technical realms today, and so they're only really fully understood by experts in that field.
"What has been will be again, what has been done will be done again; there is nothing new under the sun."
- Ecclesiastes 1:9 (circa 250 BC)
That's a lot like claiming no one says anything new because we're also using more-or-less the same words we were using last century.
I think you're missing a lot.
… it is still a little weird that the default is, like, psychotic levels of positivity, though.
So, yeah, between that, and that both your average person and someone arguing from Turing tapes can show AI doing things not in the training data, I'm going with actual data thats being "thrown around" rather than the rando blog post that argues from "by definition it can't be creative because it isn't original because it only repeats training data" Call me when they're at N=152 on a blind test.
The obvious problem is that these logic and arithmetic operations don't really need exhaustive training examples. Training the LLM is supposed to teach it a method or process by which it arrives at the answer, not just the question answer pair itself, which tends to only lead to very good approximate results. That is ok when using language, since your exact words don't really matter if you can express yourself in another way, but when it comes to algebras, accuracy is key.
It really is strange that when one wishes for a more capable LLM that one is told how stupid one is for wanting such a thing. What exactly is so wrong about giving the model a price list and an example of an order at a restaurant and then asking how much the whole order costs and expecting the correct total sum? For some strange reason I'm getting numbers that would make sense if I generated the text in a vacuum without the price list, as if they are following some existing pattern...
“Brainstorming” and “doing precise arithmetic” are wildly different tasks. It is a complete non-sequitur to ask how a suggestion about the former will help the latter.
I think people often forget the behavior of instruction following is still based on SFT training data. This is exactly "adherence to training patterns". Prompt engineering is not elixir, it's just a way to utilize the patterns seen in training data.
I’m not talking about prompt engineering when I talk about per request tuning of how closely it follows established patterns, I’m talking about inference parameters (temperature, top_p, top_k, are common for most models and ways of calling them, some others may be available, too.)