Freakonomics Radio
Freakonomics Radio

Policymaking Is Not a Science (Yet) (Ep. 405 Rebroadcast)

Why do so many promising solutions — in education, medicine, criminal justice, etc. — fail to scale up into great policy? And can a new breed of “implementation scientists” crack the code?

Featured Speakers

Freakonomics Radio + Stitcher HostJohn List Guest

Topics Discussed

Episode Summary

Executive Summary: The episode argues that evidence-based policy often fails not because research is bad, but because scaling from controlled trials to messy real-world systems is hard. Through examples in healthcare, education, and juvenile justice, John List, Dana Susskind, and colleagues explain implementation science, replication, fidelity, and contextual mismatch as key reasons interventions lose impact at scale.

Main Topics: The gap between research and policy (Priority: 5/5): The episode frames a core problem: strong academic findings often remain trapped in journals and do not translate into effective policy or practice. Scaling as the central challenge (Priority: 5/5): Many interventions work in small, controlled settings but fail when expanded because institutions, people, and environments differ from the original study conditions. Implementation science and fidelity (Priority: 5/5): The discussion introduces implementation science as the study of how programs are delivered in practice, emphasizing fidelity to the original model as a driver of success or failure. Three buckets of scaling failure (Priority: 5/5): List categorizes failures as: weak original evidence, the wrong people studied, and the wrong situation/context used at scale. Education and early childhood interventions (Priority: 4/5): Dana Susskind’s work on early language exposure and parent interaction illustrates how promising child-development programs must be redesigned for population-level reach. Juvenile justice and foster care reform (Priority: 4/5): Patty Chamberlain’s treatment foster care program shows both the potential and the difficulty of scaling evidence-based interventions across systems with conflicting rules. Replication as a policy safeguard (Priority: 5/5): The speakers argue for multiple independent replications before scaling programs, to reduce false positives and protect policymakers from costly failures.

Key Arguments: Good research is necessary but insufficient; policy success depends on whether a program can be implemented reliably in real-world settings. A major reason interventions fail is 'voltage drop'—their effects shrink when moved from ideal research conditions to broader, messier environments. Researchers often study the people most likely to benefit or comply, so results can overstate what will happen in a general population. Programs may fail not because the intervention is wrong, but because the delivery system, staffing, incentives, or institutional rules are incompatible. Replications should be rewarded institutionally; scholars should not scale interventions until results are independently reproduced multiple times. Fidelity measurement is essential because scaled programs often drift from the original design, reducing impact or increasing costs. Scaling requires humility: original programs may need adaptation, but changes must be tested rather than assumed to preserve effectiveness. Implementation science should bridge the divide between scientific entrepreneurs and institutions designed for efficiency rather than innovation.

Data Points: Americans with high blood pressure: about one-third - Introduced as an example of a widespread condition with poor treatment adherence Awareness of hypertension status: about 80% - John List describes the blood-pressure treatment cascade Blood pressure controlled: about 50% of those affected - Shows the gap between available treatment and actual outcomes Parent Academy outcome window: 3 to 6 months - List says the Chicago Heights intervention produced strong cognitive and executive-function gains in this period Parent Academy scale-up outcome: failed miserably - The UK rollout failed because no parents signed up 30 Million Words estimate: 30 million fewer words - Original estimate for the language gap between low-income and affluent children by age four Alternative estimate mentioned in controversy: 4 million - A later replication claimed a much smaller number, prompting debate over framing rather than the broader issue of interaction quality Youth prevention programs backed by research: 8% - Department of Education survey cited by List on prevention programs for youth substance use/crime Treatment Foster Care Oregon tenure: roughly 25 years - Chamberlain notes the model has endured and spread beyond Oregon Original scaling request: 15 sites - Federal officials asked Chamberlain to implement the foster-care program across multiple sites Original effect reduction at scale: a tenth or a quarter of the original result - Used to explain 'voltage drop' in implementation science

Pivotal Quotes: "Policy science not to be an oxymoron." — John List: Describing the goal of making evidence-based policymaking actually work at scale "We need to know what is the magic sauce." — John List: On identifying which elements of an intervention must be preserved for scaling "If we don't take care of it as scientists, I think everything we do can be undermined in the eyes of the policymaker and the broader public." — Aaron Trevor Brandon: Explaining why the scaling crisis matters for credibility

Implications: Listeners should expect more emphasis on replication, implementation, and fidelity before new programs are expanded. For policymakers and funders, the message is to test not just whether something works, but whether it can work broadly and sustainably.

🔓 Sign Up for Unlimited Episode Search

About Freakonomics Radio

Freakonomics co-author Stephen J. Dubner uncovers the hidden side of everything. Why is it safer to fly in an airplane than drive a car? How do we decide whom to marry? Why is the media so full of bad news? Also: things you never knew you wanted to know about wolves, bananas, pollution, search engines, and the quirks of human behavior. To get every show in the Freakonomics Radio Network without ads and a monthly bonus episode of Freakonomics Radio, start a free trial for SiriusXM Podcasts+ on...

View all episodes from Freakonomics Radio