80,000 Hours Podcast
80,000 Hours Podcast

#176 – Nathan Labenz on the final push for AGI, understanding OpenAI's leadership drama, and red-teaming frontier models

OpenAI says its mission is to build AGI — an AI system that is better than human beings at everything. Should the world trust them to do that safely? That’s the central theme of today’s episode with Nathan Labenz — entrepreneur, AI scout, and host of The Cognitive Revolution podcast. Links to learn

Featured Speakers

The 80,000 Hours team Host

Topics Discussed

Episode Summary

Executive Summary: The episode centers on Nathan LeBenz’s firsthand red-team experience with GPT-4, his evolving view of OpenAI, and what the Sam Altman firing revealed about AI governance. He argues OpenAI has become much more serious about safety, but still sees a dangerous gap between rapidly improving capabilities and slower control measures, especially as the company openly pursues AGI.

Main Topics: GPT-4 red-team experience and safety gaps (Priority: 5/5): Nathan recounts testing early GPT-4, finding it far more capable than expected but also revealing that safety controls were brittle, inconsistent, and easy to bypass with simple prompt engineering. OpenAI’s governance and the Sam Altman firing (Priority: 5/5): Rob and Nathan discuss the board’s surprise move, possible communication breakdowns, the possibility that safety concerns were broader than any single dispute, and why the episode remains opaque. OpenAI’s progress on product quality and safety commitments (Priority: 4/5): Nathan argues OpenAI has since made many smart moves: staged launches, improved refusal behavior, safety teams, forums, audits, and better public safety rhetoric. AGI versus narrow AI (Priority: 5/5): Nathan questions whether OpenAI should keep pursuing a single AGI system and suggests narrow, specialized AI systems may be safer and sufficient for many valuable use cases. Capability acceleration vs control lag (Priority: 5/5): A central theme is that AI capabilities appear to be improving much faster than alignment, monitoring, and safety techniques, creating a widening risk gap. Release timing, compute overhang, and public awareness (Priority: 4/5): Nathan becomes more sympathetic to releasing models like ChatGPT and GPT-4 earlier because it forced the world to confront the technology and start building governance capacity. Broader AI ecosystem, including Meta and open source (Priority: 4/5): The discussion ends by contrasting OpenAI’s cautious posture with Meta’s open-source strategy and the ease with which supposedly guarded models can be modified or uncensored.

Key Arguments: GPT-4 was already powerful enough to automate many knowledge-work tasks, but its safety layer was initially weak and often trivially bypassed. Helpful-harmless-honest is not a default property of powerful models; harmlessness requires separate engineering and cannot be assumed from raw capability. OpenAI has since earned more trust by launching cautiously, investing in safety, and advocating for regulation focused on frontier models. The AGI goal itself should be reexamined because OpenAI’s stated aim is to build something better than humans at almost everything before control is robust enough. A narrow-model strategy could deliver many benefits while reducing risk compared with one general system that can do everything. Releasing GPT-4 and ChatGPT may have accelerated public understanding, policy attention, and safety work faster than holding back would have. OpenAI’s board conflict is still opaque, but a breakdown in candor or trust may have mattered more than any single safety disagreement. OpenAI’s team collectively may be a more important check than the board, because employees can walk and shape strategy from within.

Data Points: OpenAI revenue in 2022: $25–30 million - Nathan says Waymark was one of OpenAI’s early customers when OpenAI was still relatively small. OpenAI revenue run rate by late 2023: ~$1.5 billion annual run rate - Used to illustrate OpenAI’s explosive commercial growth and product-market fit. Waymo autonomous miles in cited study: 3.8 million miles - Nathan cites Waymo/Swiss Re data to support accelerating proven AI deployment. Waymo body injury claims in rider-only mode: 0 claims - Compared with human driver baseline in the cited insurance study. Human driver baseline body injury claims: 1.11 claims per million miles - Comparison point in the Waymo safety study. Waymo property damage claims in rider-only mode: 0.7 claims per million miles - Compared against human driver baseline in the insurance analysis. Human driver baseline property damage claims: 3.26 claims per million miles - Comparison point in the Waymo safety study. OpenAI superalignment compute commitment: 20% of compute resources - Nathan highlights this as evidence OpenAI is taking safety seriously. GPT-4 benchmark lead: About 7–8 points on MMLU - Nathan says GPT-4 remained the top public model and still led the benchmark by a meaningful margin. MMLU score mentioned for GPT-4: ~87/100 - Nathan uses this to illustrate GPT-4’s enduring capability lead. OpenAI reporting threshold mentioned: 10^26 FLOPs - Referenced as a frontier-model threshold in U.S. policy discussion. GPT-4 training cost estimate: ~$100 million - Used to argue that scaling further may require enormous resources. Fine-tuning data needed to change refusal behavior: ~100 examples - Nathan says open-source models can be modified or uncensored with surprisingly little data. Fine-tuning cost: Under $1 to a few dollars - He contrasts current fine-tuning costs with how cheap and accessible model modification has become.

Pivotal Quotes: "I just don't feel like the control measures are anywhere close to being in place for that to be a prudent move." — Nathan LeBenz: His core objection to OpenAI’s continued push toward AGI. "What I saw was a capabilities explosion and a control kind of petering out." — Nathan LeBenz: His summary of what the red-team experience made him fear in late 2022. "I would just like to see us kind of stay in that sweet spot for a while." — Nathan LeBenz: His preference for powerful but still manageable AI systems rather than a rapid leap to AGI.

Implications: Listeners should expect AI capability gains to keep outpacing controls unless safety work scales dramatically. The episode argues for more scrutiny of AGI ambitions, more transparency from labs, and more reliance on broad internal and external accountability.

🔓 Sign Up for Unlimited Episode Search

About 80,000 Hours Podcast

View all episodes from 80,000 Hours Podcast