Episode Summary
Executive Summary: Russ Roberts and Philip Tetlock discuss Superforecasting and the challenge of evaluating predictions in politics, economics, and policy. Tetlock argues that vague punditry evades accountability, while structured forecasting tournaments show that some people, especially in teams and with training, can make meaningfully better probability judgments than experts. The conversation centers on calibration, base rates, aggregation, humility, and when better forecasts actually improve decisions.
Main Topics: Why vague punditry is hard to evaluate (Priority: 5/5): Tetlock criticizes public commentators for using imprecise language like 'distinct possibility,' which allows them to evade accountability after events unfold. Roberts presses on the cultural incentives that reward ambiguity and punish explicit probabilities. Forecasting tournaments and the IARPA project (Priority: 5/5): Tetlock explains the long-running forecasting research program he helped build, including the U.S. intelligence community tournament that tested thousands of forecasters on national-security questions and measured accuracy over time. Scoring forecasts and the problem of hindsight bias (Priority: 5/5): The discussion emphasizes why single events are hard to judge, why probability forecasts must be evaluated across many events, and how people retrospectively inflate their confidence after outcomes are known. Superforecasters: skill, training, and teamwork (Priority: 5/5): Tetlock argues that some forecasters consistently outperform others because they treat forecasting as a learnable skill, use base rates, remain open-minded, and benefit from training, collaboration, and disciplined group processes. Aggregation, extremizing, and wisdom of crowds (Priority: 4/5): The episode explains how combining forecasts can improve accuracy, especially when inputs are independent and diverse. Tetlock describes the use of statistical aggregation and extremizing to turn a group average into a better forecast. Limits of precision and policy relevance (Priority: 4/5): Roberts and Tetlock debate whether better probabilities matter for policy. Tetlock says the value depends on the domain: in finance, intelligence, and some policy choices, better calibration can matter a lot, but the gains must justify the cost. Adversarial collaboration and accountability in economics (Priority: 4/5): The conversation ends by proposing forecasting tournaments for public policy debates, such as minimum wage or quantitative easing, to force competing sides to make testable predictions and reduce post hoc self-justification.
Key Arguments: Vague language protects forecasters from being proven wrong, but it also makes their claims nearly impossible to evaluate. Forecasts should be judged probabilistically across many events, not by whether any single prediction turns out right or wrong. Hindsight bias causes people to remember themselves as having been more accurate than they actually were. Superforecasting is partly a skill: people can learn to reason from base rates, revise incrementally, and update with evidence. The best forecasts often come from combining diverse, independent judgments rather than relying on a single expert. Teams can outperform individuals when they are structured to avoid groupthink and encouraged to practice precision questioning. Forecasting is not pure noise: if it were, training, team design, and algorithmic aggregation would not reliably improve accuracy. Better probabilities do not automatically produce better decisions, but they often improve decision quality enough to matter in domains like finance, intelligence, and policy. Adversarial collaboration could make public policy debates more accountable by forcing participants to commit to scoreable predictions. Public intellectuals and experts would need outside pressure to participate in transparent prediction contests because current incentives reward ambiguity.
Data Points: Forecasting tournaments span: 30+ years - Tetlock describes forecasting research begun in the mid-1980s and continuing through the IARPA era. Tetlock's age at interview: 61 - He gives his age while explaining the chronology of his forecasting work. Initial forecasting research start: 1984 - He says he began the research after earning tenure at Berkeley. Early Soviet-era pilot studies: mid-1980s - Studies were run when the Soviet Union still existed and before Gorbachev became general secretary. IARPA tournament duration: 2011 to 2015 - The large intelligence-community forecasting tournament ran during these years and was still ongoing in some form. Forecasting questions: 500+ - The IARPA tournament covered hundreds of national-security questions across many domains. Forecasters involved: many thousands - The later tournament involved a very large number of participants across multiple teams. Forecast count: million-plus forecasts - Tetlock describes the volume of predictions generated in the IARPA project. Outperformance margin: pretty resounding margins - Tetlock says his team beat other teams by a clear margin in the first two years. Prediction-market example: 75% - Markets gave about a 75% chance that Obamacare would be overturned, but the Supreme Court upheld it 5-4. Supreme Court outcome: 5-4 - Used as an example of the danger of judging a well-calibrated forecast by a single contrary outcome. Greek exit example forecast: 0.49 vs 0.1 - Roberts uses this to discuss how two forecasters may both be below 50% yet differ in calibration. Skill-luck estimate: 70-30 - Tetlock estimates the IARPA tournament reflected roughly 70% skill and 30% luck based on regression-to-the-mean effects. Teams' advantage: about 10% - Tetlock says teams were significantly better than individuals, by roughly ten percent in early comparisons. Probability scale used: 100 points - He explains that forecasts were scored on a 100-point probability scale. Useful uncertainty resolution: 15 to 20 degrees - Tetlock estimates superforecasters can distinguish about 15-20 meaningful levels on the probability scale. Typical uncertainty resolution: 4 to 5 degrees - Most people distinguish only about 4-5 useful levels on the same scale. Base-rate example for dictators: 90%+ - Tetlock says a dictator who has already survived in power for a year or two often has a very high chance of surviving another year.
Pivotal Quotes: "They rely almost exclusively on what we call vague verbiage forecasting." — Philip Tetlock: Tetlock explains why pundits are hard to hold accountable for predictions. "The psychologists called that the hindsight bias or the I-knew-it-all-along effect." — Russ Roberts: Roberts reflects on how people revise their memories of prior beliefs after outcomes are known. "Accuracy, accuracy, accuracy." — Philip Tetlock: Tetlock summarizes the intelligence community’s criterion for evaluating forecasting teams.
Implications: The episode suggests that forecasting can improve with discipline, accountability, and aggregation, but only when predictions are made precise enough to score. For policy and media, that means less rhetoric and more testable probabilities.
About EconTalk
EconTalk: Conversations for the Curious is an award-winning weekly podcast hosted by Russ Roberts of Shalem College in Jerusalem and Stanford's Hoover Institution. The eclectic guest list includes authors, doctors, psychologists, historians, philosophers, economists, and more. Learn how the health care system really works, the serenity that comes from humility, the challenge of interpreting data, how potato chips are made, what it's like to run an upscale Manhattan restaurant, what caused the...