80,000 Hours Podcast
80,000 Hours Podcast

#81 - Ben Garfinkel on scrutinising classic AI risk arguments

80,000 Hours, along with many other members of the effective altruism movement, has argued that helping to positively shape the development of artificial intelligence may be one of the best ways to have a lasting, positive impact on the long-term future. Millions of dollars in philanthropic spending

Featured Speakers

The 80,000 Hours team HostBen Garfinkel Guest

Topics Discussed

Episode Summary

Executive Summary: Ben Garfinkel argues AI may still be highly important for the long-term future, but challenges classic AI doom stories. He says many familiar arguments rely on a "brain-in-a-box" discontinuity, weakly specified risks, and a too-clean separation between capabilities and alignment. He supports AI safety funding, but wants more rigorous, written, and detailed arguments before treating extinction-risk claims as settled.

Main Topics: Why AI could matter for the long run: Garfinkel explains the broad case that AI could transform civilization similarly to the Industrial Revolution, making it a key area for long-term-impact careers and policy. Political destabilization and war risk: The conversation covers how AI might affect great-power war, nuclear deterrence, autonomous weapons, misinformation, and broader institutional fragility. Lock-in and path dependence: They discuss how early AI governance, norms, and design choices might persist, and whether current decisions could shape future institutions or values for a long time. Classic AI risk arguments and the 'brain-in-a-box' model: Garfinkel critiques the idea that progress will stay narrow until a sudden leap to human-level AI, arguing real-world progress may be smoother and more distributed. Capabilities and alignment are entangled: He argues that making AI more capable and making it aligned are not separate steps; in practice, systems need enough alignment to be built and deployed at all. Instrumental convergence and deceptive behavior: He questions the idea that most intelligent systems will inevitably pursue power-seeking or omnicidal goals, and discusses mesa-optimization and treacherous-turn concerns. State of the evidence and communication in AI safety: Garfinkel says the field has been too reliant on a few influential books/blog posts, with too little rigorous public debate and too much confidence in weakly written arguments.

Key Arguments: AI may be worth working on because it could have world-historical effects similar to the Industrial Revolution, affecting future generations, institutions, and global power. Even if AI is transformative, that does not automatically imply individuals can steer it; some historical transformations were very important but not easily influenced in advance. Political instability is a plausible AI risk channel, especially via military destabilization, nuclear deterrence problems, autonomous weapons, and misinformation. However, AI may not be the only or even best lever for reducing war risk; hypersonic missiles, foreign policy, and leadership selection may matter more. Lock-in arguments are plausible but speculative: early laws, norms, and design choices might persist, but it is unclear how durable such effects are over centuries or millennia. The classic AI-risk picture assumes a sharp jump from narrow systems to a human-like "brain in a box" followed by rapid superintelligence, but AI may instead progress gradually across many intermediate systems. If AI development is gradual, people have more time to notice safety failures, build institutions, and use AI tools to improve oversight and alignment research. Capabilities and goals are deeply entangled in real machine learning practice: improving a system’s competence often also means shaping its objective structure, so there is no clean deadline where alignment must suddenly catch up. The instrumental-convergence thesis is weaker than it looks because most possible designs of a future technology are irrelevant; what matters is the development process, which may preferentially select benign and corrigible systems. Mesa optimization and treacherous-turn concerns are possible, but Garfinkel argues they are currently underdeveloped and should be treated as hypotheses rather than established conclusions. He thinks AI safety and governance are underfunded at the societal level, but career-switch advice should be more cautious; the strongest case is for people already well suited to AI or machine learning. He wants more public, detailed, text-based arguments from AI researchers and safety scholars so disagreements can be tested more rigorously.

Data Points: 80,000 Hours job board vacancies: 496 - The episode intro mentions the job board currently listing vacancies across AI, biosecurity, factory farming, governance, and more. Brain-in-a-box discontinuity estimate: Below 10% - Garfinkel says he puts the probability of a highly discontinuous, classic-style AI jump below 10%. Radical growth transition estimate: At least 1 in 3 - He says there is at least a one-in-three chance that, in the long run, growth becomes much faster than today. Smooth transition to faster world: Below 5% - He gives a below-5% estimate that the transition to much faster growth would itself be very abrupt. Single-year GDP growth threshold used as a proxy: 50% or above - He cites Paul Christiano’s example of a very discontinuous world as one where annual economic growth exceeds 50%. Alternative high-growth threshold: 25% - He notes that even 25% annual growth would still be historically unprecedented and evidence of major discontinuity. Order-of-magnitude change in concern: About 10x lower - Garfinkel says his credence in existential safety failure from classic AI arguments has dropped by roughly an order of magnitude. AI safety reference point for prioritization: At least five boss babies - He jokes that society is certainly spending less than this amount of money/effort on long-term AI safety and governance.

Pivotal Quotes: "I would really want the arguments to be much more sussed out and much more sort of well analyzed before I'd really feel comfortable advocating for that." — Ben Garfinkel: On whether people should switch careers into AI safety/gov without stronger arguments "I think there's this deep entanglement between the process of sort of giving it goals and the process of making it capable." — Ben Garfinkel: On why capabilities and alignment are not cleanly separable in machine learning "I think it'd be really hard to argue that they warrant, let's say, less than five boss babies worth of funding and effort." — Ben Garfinkel: His closing note that AI safety and governance are still underfunded at the societal level

Implications: Listeners should take AI seriously as a long-term issue, but not assume classic doom arguments are settled. The transcript pushes for more rigorous evidence, more gradualist thinking, and more nuanced career/policy prioritization across AI safety and governance.

🔓 Sign Up for Unlimited Episode Search

About 80,000 Hours Podcast

View all episodes from 80,000 Hours Podcast