Dwarkesh Podcast
Dwarkesh Podcast

Why I don’t think AGI is right around the corner

I’ve had a lot of discussions on my podcast where we haggle out timelines to AGI. Some guests think it’s 20 years away - others 2 years. Here’s an audio version of where my thoughts stand as of June 2025. If you want to read the original post, you can check it out here. Get full access to Dwarkesh P

Featured Speakers

Dwarkesh Patel Host

Topics Discussed

Episode Summary

Executive Summary: The episode argues AGI is not imminent because today’s models lack continual learning, robust long-horizon computer-use ability, and durable context accumulation—the traits that make humans useful employees. The host is bullish on near-term AI capabilities and expects major discontinuities once online learning is solved, but believes meaningful AGI-style transformation is more likely in 2028-2032 than in the next 1-2 years.

Main Topics: Why current LLMs fall short of real workplace usefulness (Priority: 5/5): The host says today’s models are impressive but only mediocre at practical short-horizon tasks like transcript rewriting, clip selection, and co-writing. The key limitation is that they do not improve over time like human workers do. Continual learning as the central bottleneck (Priority: 5/5): Human employees gain value by building context, learning from mistakes, and refining performance. LLMs lack a natural online learning loop, and prompt tuning or RL fine-tuning is not enough to replicate human-style skill accumulation. Limits of session memory and context compaction (Priority: 4/5): The host notes that models can get better within a session, but any gains are fragile and often lost when context is summarized or compacted. Long rolling context windows help, but text summaries are too brittle for many non-software tasks. Skepticism about near-term autonomous computer-use agents (Priority: 5/5): The host disputes forecasts that reliable agents can handle end-to-end tasks like doing taxes by next year, citing long horizons, multimodal complexity, weak training data, and the difficulty of solving agentic workflows. Reasoning progress is real, but not enough for AGI soon (Priority: 4/5): The host acknowledges that models like o3 and Gemini 2.5 show genuine reasoning and can produce impressive coding and planning behavior, but says that does not eliminate major capability gaps. Timelines and a log-normal view of AGI (Priority: 4/5): Despite skepticism about the next few years, the host still thinks the distribution of outcomes is wide and that major breakthroughs could arrive suddenly. After 2030, compute scaling may slow and algorithmic gains may become the main driver.

Key Arguments: LLMs are valuable but still only 'five out of 10' on many practical tasks, so they are not yet reliable substitutes for human workers. The biggest gap is continual learning: humans improve through practice and feedback, while LLMs largely stay fixed after deployment. System prompts and ad hoc feedback do not create the same durable, adaptive learning that human employees naturally develop. Session-level improvement exists, but it is temporary and often lost when context is compacted or summarized. Long-horizon computer-use agents are much harder than current demos suggest because they require longer rollouts, multimodal processing, and specialized data that the internet may not provide. Even simple-seeming algorithmic advances, like modern RL approaches for reasoning, took years to mature, which suggests computer-use autonomy will also take time. The host is bullish on AI’s long-term potential because once continual learning works, copies of AIs can share gains across the economy, accelerating diffusion dramatically. AGI timelines are highly uncertain, but the speaker believes the near-term chances of full transformation are lower than many optimistic forecasts imply.

Data Points: Podcast/blog date: June 3, 2025 - The narration is based on a blog post written on this date. Personal experimentation time: ~100 hours - Time the host says he spent trying to build useful LLM tools for podcast post-production. Task performance rating: 5/10 - Estimated quality of LLMs at short-horizon tasks like rewriting transcripts, picking clips, and co-writing essays. White-collar employment displaced if AI progress stalls: <25% - Host’s estimate of work that would disappear if progress stopped at current capability levels. Reliable computer-use agents forecast: By end of next year - Claim attributed to Anthropic researchers Sholto Douglas and Trenton Bricken. Computer-use taxes timeline (50/50 bet): 2028 - Host’s median forecast for an AI to handle small-business taxes end-to-end over a week-long workflow. Human-like on-the-job learning timeline (50/50 bet): 2032 - Host’s median forecast for an AI to learn as organically and effectively as a human across white-collar work. Training compute growth: 4x per year - Host says frontier training compute has been scaling at roughly this rate over the last decade. Reference to GPT-2 to GPT-4 timeline: 4 years - Used as an analogy for how long it took to go from weak to much stronger language capabilities. Reference to GPT-1 age: 7 years ago - Used to argue that seven years is a long time in AI progress terms.

Pivotal Quotes: "things take longer to happen than you think they will, and then they happen faster than you thought they could." — Rudiger Dornbush: Opening epigraph used to frame the uncertainty around AGI timelines. "The LLM baseline at many tasks might be higher than the average human's, but there's no way to give a model high-level feedback." — Host: Core claim explaining why deployment does not translate into human-like improvement over time. "It’s so economically valuable and sufficiently easy to collect data on all of these different jobs... such that... we should expect to see them automated within the next five years." — Trenton Bricken (as quoted by host): A pessimistic forecast from a prior podcast discussion that the host explicitly challenges.

Implications: Near-term AI may automate slices of work but not replace human employees without continual learning. Companies should expect useful tools, not full autonomy, until models can accumulate durable experience. Yet once that bottleneck falls, deployment could accelerate abruptly.

🔓 Sign Up for Unlimited Episode Search

About Dwarkesh Podcast

Deeply researched interviews

View all episodes from Dwarkesh Podcast