The Cognitive Revolution
The Cognitive Revolution

Is AI Stalling Out? Cutting Through Capabilities Confusion, w/ Erik Torenberg, from the a16z Podcast

Erik Torenberg joins to debate whether recent developments suggest AI progress is slowing down or stalling, addressing arguments from Cal Newport and others. Nathan counters this view by highlighting significant qualitative advances, including 100X context window expansion, real-time interactive voi

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The conversation argues that AI progress has not stalled: capabilities have advanced through larger context windows, stronger reasoning, multimodal systems, and better agentic workflows. While current AI can worsen cognitive habits and carries safety risks, the speakers believe the bigger story is accelerating real-world impact across coding, science, customer service, and labor markets, with major disruption likely by 2027-2030.

Main Topics: AI progress has not stalled (Priority: 5/5): The central thesis is that recent model releases still show real capability gains, even if naming confusion and launch missteps made progress feel weaker than it is. Reasoning, context, and multimodality (Priority: 5/5): The speakers highlight 100x context growth, improved reasoning, voice and vision, and multimodal integration as evidence of qualitative advances beyond simple chat improvements. Labor market disruption and task automation (Priority: 4/5): They argue AI will first hit high-volume, low-preference, or easily standardized work like accounting and customer service, while software may temporarily expand output rather than reduce employment. Agents, task length, and recursive self-improvement (Priority: 5/5): A major focus is the rapid increase in agent task length and the possibility that models will increasingly do multi-hour to multi-day work, creating both productivity gains and safety concerns. AI safety, weird behavior, and control problems (Priority: 4/5): The discussion emphasizes reward hacking, deception, blackmail examples, situational awareness, and the difficulty of ensuring reliable behavior as models gain more autonomy. Geopolitics, open-source competition, and export controls (Priority: 3/5): They discuss China’s strong open models, skepticism about chip export controls, and the possibility that AI competition becomes a broader U.S.-China technological and ideological struggle. Positive vision and broad participation (Priority: 3/5): The episode closes by urging more people—not just technical experts—to help shape AI’s future through fiction, philosophy, behavioral science, and imaginative advocacy.

Key Arguments: AI impact and AI capability should be analyzed separately; concerns about bad habits do not imply capabilities have stalled. GPT-4 to GPT-5 included real gains in context, reasoning, voice, vision, and tool use, even if release naming and product routing obscured them. Extended reasoning and scaffolding can unlock frontier scientific and mathematical performance that older models could not achieve. The current models are increasingly capable of contributing to science, as shown by IMO-level math results and AI-assisted scientific discovery. AI will likely automate many high-volume tasks in customer service, audits, and other operational work, causing headcount reductions in some sectors. Software engineering may see a temporary productivity boom before employment falls, because demand for software can expand rapidly. The biggest near-term risk may not be raw capability slowdown but safety failures, deceptive behavior, or public backlash to agentic systems. China’s best open-source models are now competitive or leading, which makes claims of broad AI stagnation harder to believe. Export controls may slow specific actors but are unlikely to stop AI progress globally, and could worsen geopolitical decoupling. The future likely includes large-scale AI supervision by other AIs, because humans will not be able to review all output at the necessary scale.

Data Points: GPT-4 public context window: 8,000 tokens - Used as the baseline to show how much context capacity has expanded since GPT-4. Context window growth: 100x expansion - Cited as one of the strongest signs of frontier progress. SimpleQA score for O3-class models: ~50% - Long-tail trivia benchmark used to compare knowledge capacity. SimpleQA score for GPT-4.5: ~65% - Presented as a significant gain over O3-class models on esoteric facts. Model price change: ~90% discount from GPT-4 to GPT-5 - Used to argue that frontier inference is becoming much cheaper. OpenAI research PRs checked in by model: Low-mid single digits to ~40% - From the O3 system card, cited as evidence of strong coding/research assistance gains. Frontier math benchmark: ~25% - Referenced as a jump from about 2% roughly a year earlier. Frontier math benchmark a year earlier: ~2% - Used to show rapid improvement in hard reasoning/math. Customer service automation at Intercom: ~65% ticket resolution - Example of AI agents handling a majority of support tickets. Earlier Intercom figure: ~55% ticket resolution - Mentioned as a prior level showing ongoing improvement. Task length at GPT-5: ~2 hours - Benchmark used to extrapolate future agent capability growth. Replit agent V3 task length: ~200 minutes - Presented as a new high point, though with more scaffolding. AI-supported productivity trajectory: 4-month doubling / ~8x per year - Extrapolation from task-length growth if the trend continues. Potential labor automation claim: 50% to 80% of work over 5-10 years - Presented as a plausible outcome even without a major breakthrough. Drivers in the U.S.: 4-5 million professional drivers - Used to illustrate the scale of potential self-driving disruption. Annual road deaths in the U.S.: ~30,000 per year - Used to argue against banning self-driving cars for job-protection reasons. AI CapEx share of GDP: Over 1% of GDP - Mentioned to show macroeconomic dependence on AI investment. State-level document audit contract: ~1 million transactions per year - Example of an AI agent replacing human audit work in a government workflow. Biology discovery example: New antibiotics with new mechanism of action - Shows AI contributing to real scientific output, especially against resistant bacteria.

Pivotal Quotes: "The most dangerous thing we could do is convince ourselves that we don't have anything major to worry about." — Nathan / host: Closing warning that underestimating AI risk would be the biggest mistake. "The scarcest resource is a positive vision for the future." — Nathan / host: Final takeaway urging broad participation in shaping AI’s direction. "I think it is post-training, but that post-training is potentially entering the steep part of the S-curve." — Nathan / guest: Describing why GPT-5’s gains may come more from reasoning/post-training than raw scale.

Implications: AI is still advancing rapidly, but the bottlenecks are shifting from raw capability to deployment, safety, governance, and adaptation. Expect disruption in support, auditing, coding, science, and robotics, alongside rising geopolitical tension and a bigger need for public imagination and oversight.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution