The One-Step Trap (In AI Research)(incompleteideas.net) |
The One-Step Trap (In AI Research)(incompleteideas.net) |
Since then, I have come to like temporally-abstract models more and more. Rolling out in time -- either step-by-step or many steps at once -- suffers from the tyranny of the specific. For long horizon planning with agents, I care (often only approximately) about where I can end up, and seldom about exactly when I end up there. Successor features, GVFs, Forward-Backward representations, and the like seem like they have an elegant approach for structuring thinking at a "high level", instead of generating exponentially large search trees by rolling out microscopic world models.
[1] https://arxiv.org/abs/2410.05364 (funnily, from around the same time / few months after Sutton's blog post)
I want a zoomed out picture, and to be able to fill in detail hierarchically, on demand. Instead, one-step models give you the full high-res local structure of the graph that would have to search through (with too many states and edges).
Instead, the more tokens LLMs use, the better their performance on many tasks. LLMs can self-correct, evidenced by the power of getting models to question themselves by emitting "Wait," in S1. https://arxiv.org/abs/2501.19393
What would the difference be?