Lex Fridman Podcast
Lex Fridman Podcast

#86 – David Silver: AlphaGo, AlphaZero, and Deep Reinforcement Learning

David Silver leads the reinforcement learning research group at DeepMind and was lead researcher on AlphaGo, AlphaZero and co-lead on AlphaStar, and MuZero and lot of important work in reinforcement learning. Support this podcast by signing up with these sponsors: – MasterClass: https://masterclass.

Featured Speakers

Lex Fridman HostDavid Silver Guest

Topics Discussed

Episode Summary

Executive Summary: David Silver traces his path from childhood programming to reinforcement learning, arguing that intelligence is best understood as an RL problem: an agent learning from interaction to maximize reward. He explains how AlphaGo, AlphaGo Zero, AlphaZero, and MuZero progressively removed human priors, revealing that self-play, deep learning, and search can yield creativity and superhuman performance across games and beyond.

Main Topics: Early programming and the origin of curiosity (Priority: 4/5): Silver describes first programming on a BBC Model B, early fascination with computers, and the influence of his father’s AI studies, which seeded his lifelong interest in machine intelligence. Why Go became the defining AI challenge (Priority: 5/5): He explains Go’s simple rules, massive search space, and need for intuition-based evaluation, making it a much harder AI problem than chess and a perfect testbed for reinforcement learning. Reinforcement learning as a theory of intelligence (Priority: 5/5): Silver presents RL as the formal problem of intelligence: an agent in an environment taking actions to maximize reward, with value, policy, and model as possible solution components. Deep learning and the power of scaling (Priority: 4/5): He argues that neural networks are unexpectedly effective in high dimensions, enabling large representational capacity and continuous improvement without obvious local-optimum traps. AlphaGo’s scientific and historical breakthrough (Priority: 5/5): The conversation revisits AlphaGo’s evolution from human-data-assisted deep learning to self-play, its victory over Lee Sedol, and the significance of creativity and new strategic ideas like Move 37. AlphaGo Zero, AlphaZero, and MuZero (Priority: 5/5): Silver explains the progression toward removing human expert data and even explicit rules, showing that self-play plus general learning algorithms can master Go, chess, shogi, and Atari. Broader implications: creativity, generalization, and meaning (Priority: 4/5): He connects RL to creativity, real-world applications like chemistry and quantum computing, and a layered view of goals/meaning that frames AI as another level of goal-achieving systems.

Key Arguments: Reinforcement learning captures the core structure of intelligence because intelligent behavior requires learning from interaction with an environment to maximize long-term reward. Handcrafted AI and fixed knowledge bases have ceilings and brittleness; systems must be able to learn for themselves to scale to complex real-world tasks. Go was a decisive challenge because intuition and position evaluation mattered far more than brute-force search, and humans had no reliable way to encode that intuition manually. Deep learning’s success was surprising because very large neural networks appear to avoid the local-optimum limitations people expected from low-dimensional intuition. Self-play is a powerful mechanism because it lets a system discover and correct its own errors, enabling progress from weak play to superhuman play without expert demonstrations. AlphaGo’s creativity came from discovering novel moves and patterns beyond human convention, proving that machine systems can generate genuinely new strategic ideas. Removing human priors makes systems more general and less brittle, which is essential if AI is to transfer beyond games into messy real-world domains. MuZero shows that an agent can learn a useful internal model of the world without being given explicit rules, strengthening the case for learning-based general intelligence. AlphaZero-style algorithms have already influenced domains outside games, including chemical synthesis and quantum computation, suggesting broad practical spillover. Silver believes intelligence can be understood at multiple layers, but as AI builders we need a clear objective/reward structure to make progress. Creativity is framed as repeated discovery: trying actions, seeing what works, and incorporating new successful patterns into the system’s behavior.

Data Points: Age of first programming: 7 years old - Silver wrote his first program on a BBC Model B microcomputer as a child. Game of Go board size: 19 by 19 grid - He describes the standard Go board and its simple rules. Estimated Go search space: 10^170 positions - Silver cites the astronomical scale that defeats traditional search methods. Human Go player base: ~50 million players - He notes Go’s deep cultural and global popularity, especially in East Asia. Early strongest program benchmark: Defeated by a 9-year-old with 9 handicap stones; by a computer expert with 29 handicap stones - Used to illustrate how weak Go programs still were around 2000. AlphaGo match outcome vs Lee Sedol: 4–1 - Silver predicted a 4–1 result before the match. AlphaGo Zero self-play result vs previous AlphaGo: 100–0 - Silver says AlphaGo Zero beat the prior version of AlphaGo decisively. AlphaZero chess result vs strongest computer chess program: Convincing superhuman performance - He says AlphaZero generalized without algorithm changes to chess and shogi. Training time for AlphaGo Zero: 40 days - He mentions the discovery timeline for opening patterns and Joseki during training. Number of games AlphaGo beat Lee Sedol: 4 out of 5 - The match result is discussed in detail, including the famous Move 37 game.

Pivotal Quotes: "I think it was really when I went to study at university... The only step of major significance to take was to try and recreate something akin to human intelligence." — David Silver: Explaining when he first fell in love with AI and why intelligence became his central research goal. "You have to have learning. You have to have learning. That’s the only way you’re going to be able to get a system which has sufficient knowledge in it." — David Silver: Arguing that learning is necessary to overcome the brittleness of handcrafted AI and knowledge bottlenecks. "My personal belief is that we've seen something of a turning point, where we're starting to understand that many abilities, like intuition and creativity... are actually accessible to machine intelligence as well." — David Silver: Closing reflection on the broader significance of AlphaGo and modern machine learning.

Implications: The discussion suggests that general, self-improving learning systems may be the most promising path to AGI and practical breakthroughs. For industry, the lesson is to favor scalable learning and self-play over brittle handcrafted expertise.

🔓 Sign Up for Unlimited Episode Search

About Lex Fridman Podcast

Conversations about science, technology, history, philosophy and the nature of intelligence, consciousness, love, and power. Lex is an AI researcher at MIT and beyond.

View all episodes from Lex Fridman Podcast