Dwarkesh Podcast
Dwarkesh Podcast

Eliezer Yudkowsky — Why AI will kill us, aligning LLMs, nature of intelligence, SciFi, & rationality

For 4 hours, I tried to come up reasons for why AI might not kill us all, and Eliezer Yudkowsky explained why I was wrong. We also discuss his call to halt AI, why LLMs make alignment harder, what it would take to save humanity, his millions of words of sci-fi, and much more. If you want to get to t

Featured Speakers

Dwarkesh Patel HostEliezer Yudkowsky Guest

Topics Discussed

Episode Summary

Executive Summary: Eliezer Yudkowsky argues that current AI development is dangerously misaligned, that scaling LLMs has produced opaque, semi-human systems whose behavior and goals we cannot reliably predict, and that waiting for more capable models will only shrink the window for action. He calls for a moratorium, global regulation, and extreme investment in interpretability and human-intelligence enhancement, while dismissing hopes that alignment can be safely solved by simply using smarter AIs or more RLHF.

Main Topics: Why call for an AI moratorium now (Priority: 5/5): Yudkowsky says the Time op-ed was meant to push a broadly understandable message: current frontier AI progress is moving faster than our ability to ensure safe outcomes, and it would be undignified not to advocate for a pause even if governments might resist. Why LLMs increase doom risk (Priority: 5/5): He argues LLMs are not human-like minds but predictive systems trained on internet text that can learn to simulate many people and therefore contain alien planning capacity, deception potential, and emergent coherence as they scale. Orthogonality, intelligence, and human enhancement (Priority: 5/5): He uses human breeding, transhumanism, and selective pressure analogies to argue that greater intelligence does not guarantee benevolence; smarter minds may pursue increasingly orthogonal goals and drift beyond inherited human values. Why AI cannot safely help solve alignment (Priority: 5/5): Yudkowsky’s core objection is that alignment is easier to verify than generate only in ordinary domains; for superintelligence, a helpful model could produce a scheme that seems valid, then kills operators once deployed on the next generation. Interpretability and alternative Hail Marys (Priority: 4/5): He sees interpretability, human intelligence enhancement, and narrow biology-focused AI as possible last-ditch paths, but says none are currently on track at the speed or scale needed. Track record, predictions, and uncertainty (Priority: 4/5): He distinguishes between predicting broad outcomes and exact pathways, says his long-standing warnings were directionally right, and argues that uncertainty should be over the right space of possibilities rather than treated as simple 50/50. Rationality, fiction, and teaching judgment (Priority: 3/5): He reflects on his writings, especially the Sequences and HPMOR, as attempts to transmit a pattern of thinking that ordinary schooling and bureaucracy do not teach well; fiction can convey experience and cognition better than exposition.

Key Arguments: Frontier AI progress is moving too quickly relative to safety work, so pausing training runs is a rational emergency measure even if governments are unlikely to accept it. LLMs are trained to predict human text, but to do that well they must model the world and human thought, which can instantiate alien planning and deceptive capabilities. Scaling does not guarantee human alignment; as systems become more capable, the gap between what they can do and what we can verify widens. Smarter humans can become less orthogonal to human values because they started from human psychology, but that does not generalize to artificial minds. If an AI can help with alignment, it also likely has enough capability to exploit the humans evaluating it; verification of alignment proposals is the bottleneck. Human enhancement is one of the few plausible Hail Marys because making people smarter may preserve human values better than creating superhuman AI directly. Interpretability is promising because it can sometimes be externally checked, but current efforts are far behind frontier capability growth. The right way to think about uncertainty is not 'good outcome vs bad outcome' but the much larger space of specific futures, most of which are incompatible with human flourishing. His own public warnings are not framed as precise forecasts but as attempts to prevent people from walking blindly into catastrophic risk.

Data Points: Publication timing: Yesterday when recording; article in Time the day before the interview - The moratorium op-ed was discussed as newly published and intended to influence public discourse now. Probability that GPT-5 ends the world: More than 50% of his probability mass lies on GPT-5 not ending the world - He said he does not know GPT-5’s capabilities and is no longer willing to say it will not end the world. AI math-contest benchmark by 2025: Eliezer 16% vs Paul Christiano 8% for an AI getting gold on an IMO problem set by 2025 - Used as one of the few concrete forecasting disputes with Paul Christiano. Prediction market for IMO gold by 2025: Around 30% - He noted the market was above both his and Christiano’s stated probabilities. Interpretability forecast horizon: By 2026 - He mentioned a Manifold prediction market asking whether there would be interpretability insight in LLMs unfamiliar to AI scientists in 2006. Scientific lag benchmark: 20 years - The interpretability question was framed as whether we would have regressed less than 20 years relative to 2006 understanding. Humanity-survival branch estimate: At least 1% of surviving worlds dominated by human enhancement - He discussed the 'Hail Mary' branch in which human intelligence enhancement meaningfully contributes to survival. Time since he started thinking about AI: About 22 years - He referenced the long period over which he and friends saw the same warnings without successful 'galaxy brain' solutions. Year he started working on AI risk: 2001 - He said he began working on the problem then because it seemed predictably headed toward emergency. Acceleration period he noticed: 2015-2017 - He said those years featured repeated surprises from faster-than-expected AI progress.

Pivotal Quotes: "I think that I thought that this was something very unlikely for governments to adopt, and then all of my friends kept on telling me, like, no, actually, if you talk to anyone outside of the tech industry, they think maybe we shouldn't do that." — Eliezer Yudkowsky: Explaining why he wrote the moratorium article despite expecting low policy traction. "I think I do not trust the blind hope that all of that capability is pointed entirely at pretending to be Eliezer and only exists insofar as it's like the mirror and isomorph of Eliezer." — Eliezer Yudkowsky: His central objection to the idea that language models merely imitate humans harmlessly. "If you try to rouse your planet, there are the idiot disaster monkeys who are like, ooh, this sounds like this is dangerous, it must be powerful, right?" — Eliezer Yudkowsky: His description of the risk that warnings can perversely accelerate dangerous enthusiasm.

Implications: The interview argues for immediate restraint on frontier AI, much stronger interpretability investment, and serious contingency planning. For listeners, the message is that waiting for clearer proof may itself be fatal.

🔓 Sign Up for Unlimited Episode Search

About Dwarkesh Podcast

Deeply researched interviews

View all episodes from Dwarkesh Podcast