Dwarkesh Podcast
Dwarkesh Podcast

Some thoughts on the Sutton interview

I have a much better understanding of Sutton’s perspective now. I wanted to reflect on it a bit. (00:00:00) - The steelman (00:02:42) - TLDR of my current thoughts (00:03:22) - Imitation learning is continuous with and complementary to RL (00:08:26) - Continual learning (00:10:31) - Concluding thoug

Featured Speakers

Dwarkesh Patel Host

Topics Discussed

Episode Summary

Executive Summary: The speaker analyzes Richard Sutton's 'Bitter Lesson' perspective on AI, arguing that while LLMs lack continual learning and sample efficiency, imitation learning is complementary to reinforcement learning and can serve as a crucial prior for developing AGI. They contend that human data, like fossil fuels, is a necessary intermediate step, and that techniques like test-time fine-tuning may replicate continual learning.

Main Topics: Steel Man of Richard Sutton's Position (Priority: 5/5): Sutton argues that LLMs are inefficient because they learn only during training from human data, not on the job, and lack true world models. He advocates for architectures enabling continual learning. Imitation Learning vs. Reinforcement Learning (Priority: 5/5): The speaker argues imitation learning is continuous with RL, serving as a short-horizon RL that provides a useful prior for learning from ground truth, as seen in AlphaGo and AlphaZero. Human Data as Fossil Fuels (Priority: 4/5): Analogizing pre-training data to fossil fuels, the speaker argues that human data is a crucial, non-renewable resource that enables progress toward AGI, similar to how fossil fuels enabled modern civilization. Continual Learning and LLMs (Priority: 4/5): LLMs lack continual learning, but techniques like supervised fine-tuning as a tool call or in-context learning may replicate it, as models already show flexibility within context windows. World Models in LLMs (Priority: 3/5): The speaker argues LLMs develop deep world representations despite not being trained to model action effects, challenging Sutton's definition of world models based on process rather than capability. Path to AGI: Imitation First vs. RL from Scratch (Priority: 4/5): The speaker suggests that while bootstrapping from scratch may yield better AIs eventually, imitation learning is essential for developing the first AGI, as seen with AlphaGo's superhuman performance.

Key Arguments: Imitation learning is continuous with RL; it is short-horizon RL where the episode is a token long. Human data, like fossil fuels, is a necessary intermediate step for AGI, not a dead end. LLMs develop deep world representations despite not being trained to model action effects. Continual learning may be replicated via test-time fine-tuning or in-context learning across context windows. AlphaGo's superhuman performance despite human initialization shows imitation learning is not detrimental. Cultural learning in humans is more analogous to imitation learning than RL from scratch.

Data Points: Training data equivalent: tens of thousands of years of human experience - LLMs are trained on the equivalent of tens of thousands of years of human experience. Learning rate per episode: one bit per episode - An LLM being RL'd on outcome-based rewards learns on the order of one bit per episode, which might be tens of thousands of tokens long. AlphaGo vs AlphaZero compute: AlphaZero used much more compute than AlphaGo - AlphaZero used much more compute than AlphaGo, contributing to its superior performance.

Pivotal Quotes: "Just because fossil fuels are not a renewable resource does not mean that our civilization ended up on a dead end track by using them." — Speaker: Analogizing pre-training data to fossil fuels to argue human data is a necessary intermediate step. "Imitation learning is just short horizon RL. The episode is a token long." — Speaker: Arguing that imitation learning is continuous with reinforcement learning. "If we're not allowed to call their representations a world model, then we're defining the term world model by the process that we think is necessary to build one rather than the obvious capabilities that this concept implies." — Speaker: Challenging Sutton's definition of world models based on process rather than capability.

Implications: The debate suggests that current LLM paradigms may be a necessary step toward AGI, but future systems will likely incorporate continual learning and sample efficiency. Listeners should expect hybrid approaches combining imitation and reinforcement learning.

🔓 Sign Up for Unlimited Episode Search

About Dwarkesh Podcast

Deeply researched interviews

View all episodes from Dwarkesh Podcast