Dwarkesh Podcast
Dwarkesh Podcast

AI researchers debate how close we are to recursive self-improvement

New episode with John Schulman, Beren Millidge and Charlie O’Neill. I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next. Watch on YouTube; re

Featured Speakers

Dwarkesh Patel Host

Topics Discussed

Episode Summary

Executive Summary: The discussion centers on why AI may or may not reach superintelligence quickly, with most speakers arguing the main constraints are not raw scaling but generalization, objective specification, sample efficiency, continual learning, and deployment feedback loops. They also debate whether distillation, RL, and environment design will accelerate progress, and whether future models will become continuously improving, domain-spanning remote workers or remain bottlenecked by real-world complexity and alignment.

Main Topics: Why AI might not become superintelligent by 2036 (Priority: 5/5): The panel debates the most plausible technical reasons for a slower-than-expected AI trajectory: weak generalization, inability to discover new paradigms, persistent bottlenecks in judgment, and hard limits on continual learning and self-improvement. Scaling laws, RL, and the limits of the current paradigm (Priority: 5/5): Speakers compare pretraining scaling to RL-era improvements, arguing that progress may continue in cycles of diminishing returns punctuated by new discontinuities rather than a smooth singularity. Environment design, distillation, and the role of real-world data (Priority: 5/5): Much of the conversation focuses on how labs create training environments, how distillation can copy behavior from stronger models, and how deployment data could become a major source of future capability gains. Continual learning, sample efficiency, and deployment feedback loops (Priority: 4/5): The guests debate whether models can truly learn on the job from real usage, or whether catastrophic forgetting, weak online updates, and economic incentives will keep learning fragmented and indirect. Alignment, objective specification, and human judgment (Priority: 4/5): A recurring claim is that humans may remain in the loop the longest for defining objectives, model behavior, and alignment rules, even if technical execution becomes highly automated. Forecasts for AI researchers and remote workers (Priority: 4/5): The panel gives timelines for AI research acceleration, general white-collar automation, and eventual cross-domain dominance, with predictions ranging from 1-3 years for strong remote-worker-like systems to 3-10 years for broader superhuman capability.

Key Arguments: The main technical failure mode for superintelligence is likely not raw compute limits, but failure to achieve robust generalization, meta-learning, and continual learning across shifting tasks. Current models already show repeated cycles of seeming impressive, then feeling dumb after use; this may continue if models keep bottlenecking on judgment and self-checking. RL has helped more than expected because mid-training and high-signal environments warm-start models, and because small policy tweaks can cause large functional changes. Distillation can efficiently copy capabilities once a teacher exists, but it depends heavily on realistic prompt distributions and can lag on broad, messy real-world behavior. Deployment data is increasingly valuable; future progress may come from turning real usage into training signals via distillation, online learning, or modular updates. Humans may remain necessary longest for specifying what the model should do, especially in alignment, constitutions, and behavior design. The path to AI that automates AI R&D likely involves multi-step research environments plus human feedback, not a single clean self-play loop. Sample inefficiency and catastrophic forgetting may make true continuous learning difficult, limiting how fast models can absorb new real-world experience. AIs may become strong at cumulative tasks like RSI faster than at non-stationary real-world work, because RSI is more objective-stable and easier to formalize. The panel generally expects major acceleration if models can run multiple experimental loops autonomously, but they disagree on whether that becomes enough for ASI quickly.

Data Points: Timeline for full general remote-worker-like AI: ~1-3 years - Predictions for an AI that can do month-long white-collar work with computer use and some human interaction Timeline for 10x productivity uplift for AI researchers: ~2 years (with some saying 5-10 years for full uplift) - Estimates for AIs that significantly speed up AI research and experimentation Timeline for AI that dominates all human experts across computer-based work: ~3-4 years (one estimate), or ~5-10 years when including longer learning and broader domains - Predictions for broad superhuman cognitive performance Reported pretraining data efficiency gain: ~9x - A cited investigation comparing recipe and data changes from 2019 onward Reported architecture efficiency gain: ~3x - Same investigation; architecture improvements contributed less than data at small scale Claimed cumulative improvement from 2019 scaling estimates: ~27x observed vs ~2000x implied by 3x/year over 7 years - Used to argue that much of the missing gain likely comes from post-training or other factors Parameter scaling guess for frontier models: Active parameters may plateau near current hundreds of billions to low trillions for a few years - Discussion of inference efficiency and RL constraints Deployment to training cadence at some companies: Every ~5 hours in one example - Cursor reportedly deployed new versions frequently if they improved benchmark performance Example performance result: A fine-tuned 1930s-era Talkie model beat Claude 3 Opus on SWE-bench - Used to show how much expert behavior can be copied with enough task data Reported Horizon/generalization trend: Doubling every ~3 months - Referenced via an EdgeBench-style claim that models can work longer for extended horizons

Pivotal Quotes: "Alignment is the final job." — John Shulman: On what humans are likely to keep doing the longest even as AI takes over technical work "The only hope really is if deep learning just can't get us to an AI which is at least can dominate human research and human development." — John Shulman: On the main technical reason a superintelligence scenario might fail "The thing that just did this originally was I was saying, isn't it weird how Sonnet 5 and Opus 5 are almost objectively worse models than GLM 5.3, Qimi K3..." — Host: A prompt about why frontier models may lag distillable competitors depending on prompt distribution and deployment data

Implications: The panel expects fast but uneven progress: strong gains in coding, research, and white-collar automation, with the biggest uncertainty around learning-from-deployment, objective design, and continual adaptation. If those are solved, capability growth could accelerate sharply; if not, AI may keep improving in bursts without full superintelligence.

🔓 Sign Up for Unlimited Episode Search

About Dwarkesh Podcast

Deeply researched interviews

View all episodes from Dwarkesh Podcast