What is the status on continual learning for LLMs? I belive if this is solved, agi is very close. So id like to know if its solvable |
What is the status on continual learning for LLMs? I belive if this is solved, agi is very close. So id like to know if its solvable |
Even if we had an algorithm to incrementally update weights without catastrophic forgetting (I don't believe we do - but doesn't seem like such a tough problem), then this implies that everyone has their own personalized model, else if you combine all these updates there is no data privacy. Cloud serving also really depends on everyone using the same model so that you can batch requests and not reload weights for each user.
The much more achievable goal, without needing to upend the whole serving model, would be just to implement continual "episodic" memorization (text-only maybe), but even with compaction/consolidation the next question would be how do you retrieve these external memories and get the necessary chunks into context when needed. Some sort of vector store, perhaps? How do you avoid vendor lock-in - perhaps have agents store/retrieve vector-store context chunks from a vendor-agnostic cloud store?
So, even the most basic form of continual-learning-like enhancement becomes tricky. What you'll first see is presumably just enhancements of the vendor-specific memory mechanisms that are already available.
What I would consider as true continual learning, close to what an animal or human does, would require much more extensive architectural and deployment/business model changes - since then we're really talking about more than just an update/recall problem, but rather the whole autonomous agentic loop of predicting/acting/failing/learning/etc with innate traits like curiosity (prediction failure) and boredom to make sure the agent is exposed to learning situations in the first place. At this point you are building an artificial brain, not just an LLM. Some companies such as Google/DeepMind have a more ambitious definition of AGI (more than just an LLM) that is perhaps a step in this direction.
Even with this sort of animal-like continual exploration/learning, you still wouldn't have something that is human level, but at least much closer in terms of ability to learn.
The difference with something completely new at inference time, that was not in the training data, or not used by RL training, is that while it will be able to use it to some extent, it has not been taught via RL how to best use in a reasoning chain.
The problem is that LLMs don't really have the generic ability to reason, so they instead need to fake it by either:
1) Fine tuning on reasoning data (very specific)
2) RL training (more generalizable, especially for math/coding)
3) Prompting that encourages "keep on going" "tree of thoughts" exploration where they may discover a reasoning chain largely by luck
Demis Hassabis has talked about combining LLMs with search (cf systems like AlphaGo) which sounds like it might be useful for math research.
> also i think that it is reasoning exactly like us humans do. Indistinguishable
Yes, it is copying human reasoning so it will appear the same, but the difference is when you don't know what to do/try next - when you are trying to solve a problem that you have never solved before and don't know from experience what to try next. This is when having real/generic ability to reason, not just "reason from memory" matters. This is when things like human curiosity are useful : "I wonder what happens if I try this ..."
I think for true, especially open-ended, research you really also need things like curiosity, but perhaps a lot can still be done with the same sort of prompting that got the Jacobian result - just tell the model to keep on going, investigate anything that seems interesting/unexplained, etc?
I don't know what it would take to have an AI that might make a brand new discovery/theory like this... part of it would be the mechanics of continual learning, curiosity, etc, but maybe also part education/mentorship?. How/when does a human mathematician, tackling a problem and getting nowhere (or maybe via some other motivation), end up inventing a whole new framework/paradigm to make progress ?!
I think it would be very surprising to see today's AI (LLMs) do this, because there is just too much missing. LLMs were designed only to be next token predictors - i.e. copying/automation machines. There are trained to predict, hence copy, what they have seen before, not to innovate (other than by combining what they have already seen). Humans may learn and process language in a way similar to an LLM, based on prediction, and our brains are certainly evolved for prediction, but there is a lot more to us than that. At least half of our cortex is dedicated to feedback and learning - what to do when our predictions are wrong. LLMs don't have this. When an LLM prediction "fails" it hallucinates. When a human prediction fails other mechanisms kick in (curiosity, boredom, frustration), some purely cognitive, some emotional, and we potentially explore, innovate and overcome.
Presumably if you try the RAG (vector store) approach, then one necessary component would be for the system to continually review what it has already learnt and consolidate/expand that by looking for generalizations, exceptions, contradictions, implications, etc. In fact maybe this would be the core of the system, with subagents assigned to each of these (and more) tasks. You would just "seed" the memory with the task you wanted it to work on, some hints on direction, difficulties, etc, then let it get to work. This would be rather like a GOFAI "blackboard" system where a bunch of loosely collaborating agents make progress on a problem by using a shared blackboard to get new work and post results (which then become inputs to other agents).