Freakonomics Radio
Freakonomics Radio

494. Why Do Most Ideas Fail to Scale?

In a new book called "The Voltage Effect," the economist John List — who has already revolutionized how his profession does research — is trying to start a scaling revolution. In this installment of the Freakonomics Radio Book Club, List teaches us how to avoid false positives, how to know

Featured Speakers

Freakonomics Radio + Stitcher Host

Topics Discussed

Episode Summary

Executive Summary: The episode centers on John List’s "The Voltage Effect," arguing that many ideas fail not because they’re bad, but because they don’t scale. Using Kmart’s Blue Light Special, education, Uber, Facebook, tax enforcement, UBI, and quitting as examples, List urges researchers and leaders to design for real-world constraints, test for false positives, and think marginally and systemically before expanding.

Main Topics: The Blue Light Special as a scaling cautionary tale (Priority: 5/5): Kmart’s centralized Blue Light Special lost its local flexibility, showing how a successful idea can be damaged when scaled without preserving the conditions that made it work. Policy-based evidence vs. evidence-based policy (Priority: 5/5): List argues researchers should design studies with scaling constraints in mind from the outset, rather than proving small pilots and hoping they translate later. Voltage drops and false positives (Priority: 5/5): Many promising interventions show strong results in pilots but weaken or reverse at scale due to general equilibrium effects, bad measurement, or unique pilot conditions. Incentives, marginal thinking, and scalable design (Priority: 4/5): Successful scale often depends on low-cost incentives and decisions based on marginal effects rather than averages, as shown by the Dominican tax campaign and Uber/Lyft examples. Human systems, diversity, and power imbalances (Priority: 4/5): Programs fail when they assume people behave like rational automatons; scaling requires diversity, empathy, and awareness of how institutional power dynamics affect outcomes. Knowing when to quit (Priority: 4/5): List frames quitting as an optimization problem: people should compare themselves to peers and weigh opportunity costs, not just persevere for its own sake. Science, replication, and trust (Priority: 4/5): List criticizes shaky findings and fraud in social science, calling for rapid replication, access to data, and safeguards like Benford’s law to improve credibility.

Key Arguments: A great idea does not automatically become a great scalable idea; scale exposes hidden constraints and assumptions. Researchers often optimize pilot studies for success rather than for generalization, creating misleading evidence for policymakers. The key question is not whether something works in a small test, but whether it still works when implemented by ordinary people, in ordinary conditions, at larger scale. Many scaling failures are actually failures of common sense: ignoring incentives, local variation, and the behavior of marginal users. Human-centered systems are especially hard to scale because people are not perfectly rational and institutions often rely on exceptional individuals. Diversity improves scaling by broadening the set of perspectives and experiences inside organizations, which can improve both recruitment and problem-solving. Some interventions fail because they are false positives rather than true effects; replication and data transparency are essential before expansion. Quitting can be rational when comparative advantage disappears or opportunity costs rise; persistence is not always virtuous. Network effects can make platforms like Facebook valuable, but they also amplify harms when misinformation or bad content spreads. Universal basic income may look promising in pilots, but scaling could trigger labor-market and community spillovers that small tests cannot reveal.

Data Points: Kmart locations at peak: more than 2,300 - The transcript notes Kmart once had over 2,300 U.S. locations. Tax-compliance intervention messages: more than 80,000 - List and colleagues sent over 80,000 messages in the Dominican Republic tax experiment. Additional tax revenue: $100 million - The Dominican Republic messaging campaign brought in an extra $100 million. Facebook market cap: around a trillion dollars - Used to illustrate how a scaled platform creates both value and risk. Facebook users: nearly 3 billion - Shows the magnitude of Facebook’s scale and network effects. Uber growth under Travis Kalanick: $66 billion in seven years or so - Cited as evidence of Kalanick’s confidence and Uber’s rapid scaling. Pilot wage increase test: 5% of drivers - Lyft wage experiment in which only 5% of drivers received a raise in the pilot. Pilot outcome: everyone’s happy; more money per hour - In the small test, driver wage increases appeared successful. Scaled wage outcome: about the same - When scaled up, the wage increase was offset by market equilibrium effects. Tax-message split: half and half - Dominican recipients were evenly divided between jail-threat and public-exposure messages.

Pivotal Quotes: "Most of us think that scalable ideas have some silver bullet feature, some quality that bestows a can't-miss appeal. That kind of thinking is fundamentally wrong." — John List: Read from The Voltage Effect, summarizing his core thesis about scaling. "We need to move from a mentality of creating evidence-based policy to one of producing policy-based evidence." — John List: Explaining how research should be designed with scale in mind from the beginning. "Humans just don't scale." — John List: A concise explanation for why systems centered on behavior often fail when expanded.

Implications: For policymakers, founders, and nonprofits, the lesson is to test for scalability, not just pilot success. Build for marginal effects, spillovers, and human behavior; replicate fast; and be willing to quit when the real-world math stops working.

🔓 Sign Up for Unlimited Episode Search

About Freakonomics Radio

Freakonomics co-author Stephen J. Dubner uncovers the hidden side of everything. Why is it safer to fly in an airplane than drive a car? How do we decide whom to marry? Why is the media so full of bad news? Also: things you never knew you wanted to know about wolves, bananas, pollution, search engines, and the quirks of human behavior. To get every show in the Freakonomics Radio Network without ads and a monthly bonus episode of Freakonomics Radio, start a free trial for SiriusXM Podcasts+ on...

View all episodes from Freakonomics Radio