Episode Summary
Executive Summary: Nathan LeBenz shares his insider experience as a GPT-4 red teamer, revealing initial safety lapses at OpenAI that later improved. He discusses the OpenAI board drama, the divergence between AI capabilities and control, and the critical role of employees as a check on AGI development. The conversation explores the tension between narrow AI and AGI, the inevitability of progress due to compute overhangs, and the need for continued questioning of safety measures.
Main Topics: Red Team Experience and Safety Concerns (Priority: 5/5): Nathan describes his two-month red teaming of GPT-4, finding the safety edition easily jailbroken and the team's lack of urgency. He escalated concerns to the board, leading to his removal from the program. OpenAI Board Drama and Sam Altman's Firing (Priority: 5/5): Analysis of the November 2023 board decision to fire Sam Altman, the lack of public explanation, and the staff's threat to leave. Nathan speculates on board dynamics and Ilya Sutskever's role. Capabilities vs. Control Divergence (Priority: 4/5): The rapid improvement in AI capabilities (e.g., GPT-4 outperforming humans on medical tasks) outpaces safety measures, with spear phishing still working on GPT-4. Nathan argues this divergence is a key risk. Employee Power and Responsibility (Priority: 4/5): Nathan emphasizes that OpenAI employees hold significant power to halt or redirect development, as demonstrated by the mass threat to leave. He urges them to question the AGI focus. Narrow AI vs. AGI Debate (Priority: 4/5): Discussion of whether to pursue general superhuman AI or narrow, specialized models. Nathan advocates for narrow AI as safer and suggests questioning the assumption that AGI is the only goal. AI Scouting and Understanding Capabilities (Priority: 3/5): Nathan's concept of 'AI scouting'—comprehensively tracking AI developments to avoid blind spots. He highlights the difficulty of keeping up with exponential progress. Geopolitical and Regulatory Context (Priority: 3/5): Sam Altman's stance against the 'China will do it anyway' narrative, the Biden executive order, and the need for sensible regulation focused on frontier models.
Key Arguments: OpenAI's safety efforts were initially inadequate but improved over time with strategic releases (ChatGPT with 3.5 first) and commitments like the superalignment team. The board's failure to communicate reasons for firing Sam Altman undermined trust, and the staff's collective power proved to be the real check on leadership. AI capabilities are advancing faster than control measures, as shown by persistent vulnerabilities like spear phishing on GPT-4. Narrow AI models can be extremely useful and safer than AGI; the pursuit of AGI should be questioned rather than assumed. Compute and data overhangs make AI progress inevitable; early releases help society adapt and learn about risks. Employees at frontier labs have a moral responsibility to question the direction of development and be willing to walk away if necessary.
Data Points: GPT-4 MMLU score: 87% - GPT-4 remains #1 on MMLU benchmark, 7-8 points ahead of next best models. OpenAI revenue growth: $25-30M (2022) to $1.5B run rate (2023) - Nearly two orders of magnitude growth in one year. Waymo injury claims: 0 per million miles - In 3.8 million rider-only miles, zero bodily injury claims vs. human baseline of 1.11 per million miles. GPT-4V medical performance: Outperformed humans overall - On 934 NEJM medical image cases, GPT-4V exceeded human respondents except in radiology where it matched. Red team engagement: Nathan made half of all posts - Low engagement from other red teamers; Nathan was the most active participant. Fine-tuning examples to remove refusal: ~100 examples - With just 100 examples and under a dollar, one can fine-tune an open-source model to remove safety constraints. Superalignment compute commitment: 20% of compute - OpenAI committed 20% of its compute to the superalignment team over four years.
Pivotal Quotes: "I just don't feel like the control measures are anywhere close to being in place for that to be a prudent move." — Nathan LeBenz: Discussing OpenAI's stated goal of building AGI that is more capable than humans at everything. "The team is the check that is really going to matter. Nobody else, you know, the board can't check you because you guys can just all walk." — Nathan LeBenz: Emphasizing the power of OpenAI employees to halt or redirect development. "We shouldn't trust any individual person here." — Sam Altman (quoted by Nathan): Sam Altman's comment after the board drama, acknowledging the need for distributed governance.
Implications: The podcast underscores the fragility of current AI governance structures and the critical role of informed employees. It suggests that society must accelerate safety research and consider narrow AI as a safer path. Listeners are urged to stay engaged and question assumptions about AGI inevitability.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co