Dwarkesh Podcast
Dwarkesh Podcast

Will scaling work? [Narration]

This is a narration of my blog post, Will scaling work?. You read the full post here: https://www.dwarkeshpatel.com/p/will-scaling-work Listen on Apple Podcasts, Spotify, or any other podcast platform. Follow me on Twitter for updates on future posts and episodes. Get full access to Dwarkesh Podcast

Featured Speakers

Dwarkesh Patel Host

Topics Discussed

Episode Summary

Executive Summary: This podcast analyzes whether scaling large language models (LLMs) will lead to AGI. It presents a debate between a 'Believer' and a 'Skeptic' on key issues like data availability, the effectiveness of self-play/synthetic data, benchmark validity, and whether LLMs genuinely understand the world. The author gives a 70% probability that scaling plus algorithmic and hardware advances will yield AGI by 2040, but acknowledges significant uncertainties, especially if synthetic data methods fail.

Main Topics: Data Bottleneck and Synthetic Data (Priority: 5/5): The Skeptic argues we'll soon run out of high-quality language data, needing 5 orders of magnitude (100,000x) more data than available. The Believer counters that synthetic data from self-play can work, drawing analogies to human evolution and citing researchers' confidence. Validity of Scaling Laws and Benchmarks (Priority: 5/5): The Believer notes performance scales consistently over 8 orders of magnitude and that GPT-4's performance was predictable from smaller models. The Skeptic argues benchmarks like MMLU test memorization, not intelligence, and that models fail on long-horizon tasks (e.g., SWEBench score 1.7% for GPT-4). Model Understanding and Compression (Priority: 4/5): The Believer claims next-token prediction forces models to learn world regularities, leading to understanding. The Skeptic contends LLMs merely compress data via gradient descent, which is not equivalent to human-like intelligence or insight. Economic and Historical Precedents (Priority: 3/5): The Believer compares potential AI investment to historical spending on railways (7% of UK GDP in 1847) and telecom (almost $1 trillion), suggesting society will fund massive scale-ups like a GPT-8 model using 1% of world GDP. In-Context Learning and Grokking (Priority: 3/5): The Believer argues larger models naturally develop more efficient meta-learning (grokking). The Skeptic counters that grokking is far less efficient than human insight learning and not possible with current training methods. Comparison to Primate Brain Evolution (Priority: 2/5): The Believer uses neuroscience research to argue that primate brains (like transformers) have a scalable architecture that evolution 'discovered,' analogous to how transformers scale better than LSTMs.

Key Arguments: Scaling has worked consistently for eight orders of magnitude; this trend should continue. Synthetic data from self-play can work, with base models already getting '1 in 100' correct answers. Benchmarks like MMLU are poor measures of real intelligence; they only test memorization. Models are plateauing on common benchmarks (e.g., Gemini vs. GPT-4) and perform poorly on long-horizon tasks (SWEBench 1.7% for GPT-4). There is no evidence that RL or self-play increases underlying model abilities. A new architecture would need a bigger jump in sample efficiency than LSTMs to transformers (invented in the 1990s). Society is willing to spend large fractions of GDP on general-purpose technologies (e.g., 7% on railways in 1847). Insight-based learning requires a 'drag and drop' to new concepts, impossible with gradient descent.

Data Points: Compute needed for human-level AI: 1E35 FLOPs - Estimate for an AI reliable enough to write a scientific paper. Data shortfall: 5 orders of magnitude (100,000x) - Amount of high-quality language data we are short by. Additional compute for human-level thought: 9 orders of magnitude - On top of the largest models today, even with self-play. GPT-4 training cost: 0.03% of Microsoft's yearly revenues - Cost to train GPT-4 from a big scrape of the internet. Performance on SWEBench (GPT-4): 1.7% - Autonomously completing pull requests. Performance on SWEBench (Claude 2): 4.8% - Slightly more impressive but still low. Scale-up from GPT-3 to GPT-4: 100x - A relatively small increase compared to potential future scale-ups. GDP spent on railways in 1847 (UK): 7% - Historical precedent for investing in general-purpose technologies. US telecom investment after 1996 Act: $500 billion (almost $1 trillion today) - Another historical precedent for large-scale investment.

Pivotal Quotes: "My guess is that this [data] will not be a blocker. Maybe it would be better if it was, but it won't be." — Dario Amodei (quoted by Believer): On whether lack of data will prevent scaling towards AGI. "If you can keep scaling LLMs and get better and more general performance as a result, then there's reason to expect powerful AIs by 2040 or much sooner." — Dwarkesh Patel (narrator): Setting up the core premise of the debate between Believer and Skeptic. "If these models can't get anywhere close to human level performance with the data a human would see in twenty thousand years, we should entertain the possibility that two billion years' worth of data also wouldn't do the trick." — Skeptic: Argument against the scalability of LLMs to human-level intelligence.

Implications: The analysis suggests AGI by 2040 is plausible (70% chance) if scaling and synthetic data succeed, but failure could mean a much longer path. Listeners and investors should watch for progress on self-play and long-horizon task benchmarks as key indicators.

🔓 Sign Up for Unlimited Episode Search

About Dwarkesh Podcast

Deeply researched interviews

View all episodes from Dwarkesh Podcast