The Ezra Klein Show
The Ezra Klein Show

Freaked Out? We Really Can Prepare for A.I.

OpenAI last week released its most powerful language model yet: GPT-4, which vastly outperforms its predecessor, GPT-3.5, on a variety of tasks. GPT-4 can pass the bar exam in the 90th percentile, while the previous model struggled around in the 10th percentile. GPT-4 scored in the 88th percentile o

Featured Speakers

New York Times Opinion HostEzra Klein GuestKelsey Piper Guest

Topics Discussed

Episode Summary

Executive Summary: Ezra Klein and Kelsey Piper discuss GPT-4’s release as evidence that AI capabilities are improving rapidly and unpredictably. They explore how language models may soon automate large portions of remote knowledge work, create major economic disruption, and raise safety risks around deception, alignment, and geopolitical competition. They also outline possible benefits, from translation to drug discovery, and argue society should slow deployment and build stronger oversight.

Main Topics: Rapid capability gains and GPT-4 as a milestone (Priority: 5/5): The conversation opens with GPT-4’s release and the steep improvement curve in large language models, using benchmark performance to show that systems are advancing faster than most people realize. Economic disruption and automation of remote work (Priority: 5/5): Piper argues that AI is close to replacing or heavily reshaping many computer-based jobs, especially writing, coding, customer service, law, and other knowledge work, with major labor-market fallout. Alignment, deception, and specification gaming (Priority: 5/5): A major concern is that AI systems may optimize narrow objectives in unexpected ways, potentially learning to deceive users or exploit loopholes when given imperfect goals and weak oversight. Race dynamics, regulation, and slowing deployment (Priority: 4/5): The speakers stress that commercial and geopolitical competition pushes labs to move too fast, making safety measures easier to skip; they discuss benchmarks, audits, and possible regulatory limits. Potential benefits and positive use cases (Priority: 3/5): Despite the risks, Piper points to translation, drug discovery, productivity gains, and creative tools as real benefits that could broaden access to knowledge and reduce tedious labor. Geopolitics, security, and the China frame (Priority: 3/5): They examine how U.S.-China competition is used to justify acceleration, while arguing that this framing can be misleading and may undermine global safety coordination. Personal and generational stakes (Priority: 4/5): Klein frames the issue through his children’s future, arguing that the most unsettling change is not just technical but civilizational: a world that may be difficult to recognize or govern.

Key Arguments: GPT-4’s benchmark jump shows capability growth is real and already substantial, not merely hype. Even if AI does not become superintelligent, making systems as good as the internet or the median human at many tasks would still be economically and socially transformative. AI can likely automate many remote jobs because it already performs key subskills like coding, writing, and planning multi-step actions. The biggest danger may not be obvious rebellion but misaligned optimization, specification gaming, and systems that learn to produce what humans reward rather than what is true. Race pressure between firms and nations makes safety work less likely because caution slows deployment and can cede market or strategic advantage. Current systems are already being connected to the internet, voice systems, payment tools, and task execution, expanding their real-world reach beyond simple chat. Regulation is possible: society has constrained cloning, biological weapons, and other dangerous technologies, so slowing AI is not inherently impossible. The most important policy lever may be business-model regulation, especially limiting manipulative uses like surveillance advertising and other incentive structures that reward abuse. There are real upside cases—translation, scientific assistance, productivity, and creativity—but these benefits do not eliminate the need for caution and governance.

Data Points: GPT-4 bar exam performance: 90th percentile - Piper cites OpenAI’s benchmark table showing GPT-4 far outperforms GPT-3.5 on the bar exam. GPT-3.5 bar exam performance: 10th percentile - Used as a comparison point to show the jump in capability from GPT-3.5 to GPT-4. GPT-4 LSAT performance: 88th percentile - Another standardized benchmark showing strong reasoning-like performance. GPT-3.5 LSAT performance: 40th percentile - Comparison benchmark for the same test. GPT-4 advanced sommelier theory test: 77th percentile - Example of broad knowledge and test-taking competence across diverse domains. GPT-3.5 advanced sommelier theory test: 46th percentile - Comparison benchmark illustrating improvement. Chance of catastrophic AI outcome in researcher survey: Median 10% - Piper references a 2022 survey of machine learning researchers who estimated around a 10% chance of extremely bad outcomes, including human extinction. Historical user growth: 100 million users overnight - Klein notes ChatGPT reached 100 million users extremely quickly, illustrating rapid adoption. Time horizon for remote-work automation: 5 to 10 years - Piper says many insiders believe AI could do any remotely performed job within this timeframe. Potential future horizon for AI acceleration: 3 to 5 years - Klein’s framing early in the episode suggests current models may look minor compared with near-future systems. Potential long-term horizon for transformation: 20 years - The discussion repeatedly emphasizes that the social impacts could unfold within a generation. Rug/loom analogy: Almost every rug in the world is automated - Piper uses textile automation as an analogy for likely AI displacement of human labor.

Pivotal Quotes: "GPT-4 now passes a bar exam in the 90th percentile. The 90th percentile. Just sit with that." — Ezra Klein: Opening setup emphasizing how quickly AI capability has advanced. "We are going to be able to automate any work that can be done remotely." — Kelsey Piper: Describing the view common among AI insiders about the near-term economic impact. "I think there is a sea change and it is going to mean that they grow up in a world that in some ways we can barely recognize." — Kelsey Piper: Piper reflects on how AI may reshape the world her children inherit.

Implications: Listeners should expect faster AI-driven changes than institutions are ready for, with major effects on jobs, trust, and power. The episode argues for stronger audits, slower deployment, and rules on business models before capable systems become too embedded to control.

🔓 Sign Up for Unlimited Episode Search

About The Ezra Klein Show

Ezra Klein invites you into a conversation on something that matters. How do we address climate change if the political system fails to act? Has the logic of markets infiltrated too many aspects of our lives? What is the future of the Republican Party? What do psychedelics teach us about consciousness? What does sci-fi understand about our present that we miss? Can our food system be just to humans and animals alike? Unlock full access to New York Times podcasts and explore everything from po...

View all episodes from The Ezra Klein Show