Episode Summary
Executive Summary: John List explains the “voltage effect”: ideas that look promising in small tests often lose impact when scaled because of false positives, audience mismatch, non-representative settings, spillovers, and supply-side constraints. Using examples from a Chicago pre-K, Uber tipping, surge pricing, and Lyft ad spending, he argues that successful scaling requires designing experiments with scale in mind from the start and thinking on the margin.
Main Topics: The Voltage Effect and Why Ideas Fail at Scale (Priority: 5/5): List defines the voltage effect as the drop in impact when a program moves from a small pilot to a larger rollout. He identifies five causes: false positives, wrong target audience, non-representative situations, spillovers, and supply-side constraints. Designing Experiments for Scale (Priority: 5/5): Roberts and List discuss how standard A/B tests often measure efficacy rather than scalability. List argues that experiments should include “scale tests” from the beginning so results reflect real-world implementation conditions. Chicago Heights Pre-K and the Difference Between Horizontal and Vertical Scaling (Priority: 5/5): List uses his pre-K experiment to show that hiring the best possible teachers can maximize results in one site, but a scalable model must use teachers similar to what typical districts can hire. He distinguishes horizontal replication across cities from vertical expansion within one labor market. Uber, Lyft, and Tipping as a Market Design Experiment (Priority: 5/5): List recounts how Uber adopted in-app tipping after the #DeleteUber backlash and how the change affected driver supply and wages. The discussion shows how a policy can improve behavior in a pilot but wash out once broadly adopted. Surge Pricing and Supply Response (Priority: 4/5): Roberts and List debate surge pricing. List argues that while surge clearly reduces demand, its ability to instantly pull many more drivers from couches is much smaller than people assume; predictable surges are easier to manage than unanticipated ones. Marginal Thinking in Decision-Making (Priority: 5/5): List’s Lyft example shows why average data can mislead decisions. He advocates using the most recent, smallest relevant slice of data to estimate marginal returns, rather than relying on large historical averages. The Long-Run Economics of Ride-Hailing (Priority: 4/5): List suggests Uber/Lyft’s current model may be only partially viable and that the largest future profits may come from autonomous vehicles, freight, or other scaled services where capital ownership captures the rents.
Key Arguments: Small pilots often overstate effectiveness because they are run in unusually favorable conditions; true scaling requires testing under realistic constraints. A/B tests are often efficacy tests, not scale tests, so policymakers and firms can be misled when they generalize results. Teacher quality in the pilot matters for the pilot result, but if scaling depends on hiring thousands of similar teachers, the model may fail due to labor supply constraints. Horizontal scaling (replicating across markets) is different from vertical scaling (expanding deeply within one market), and an intervention can succeed at one but not the other. Uber tipping initially changed driver behavior and brought some drivers back, but once widely adopted it did not increase hourly wages because labor supply adjusted. Ride-hailing ratings and tips are influenced more by rider fixed effects and social norms than by observable trip quality alone. Surge pricing mostly works by reducing demand and reallocating timing; its short-run supply response is smaller than many assume. Average cost data can hide the true marginal cost of acquiring the next user or driver, so decisions should be based on the newest, thinnest relevant slice of evidence. The biggest profits in ride-hailing may accrue to whoever controls safe, scalable autonomy rather than to the current dispatch platform. Market design can outperform intuition: pricing, ratings, and app structure create trust and allocation mechanisms that substitute for old taxi systems.
Data Points: Date of episode: June 23, 2022 - Podcast introduction and interview framing Chicago Heights high school completion rate: 480 graduates / 1,000 starters - Baseline described for the district before intervention Chicago Heights dropout rate: 520 dropouts / 1,000 starters - Baseline described for the district before intervention Additional high school graduates from intervention: 42 more per class - List’s estimate of the program’s effect in Chicago Heights Uber market share at Lyft’s start: ~5% to 10% Lyft market share - List describes Lyft as a much smaller rival when he was at Uber Lyft market share later: 30% to 35% across North America - List notes Lyft’s later scale after the Uber backlash period Date of #DeleteUber event: January 27, 2017 - Trump immigration executive order and taxi strike discussion Tipping rollout effect size: About 10% to 15% of trips tipped - List describes the app design and resulting tipping rate Always-tip users: 1% of people tip on every trip - Shows tipping is not universal even after in-app adoption Never-tip users: 3 out of 5 people never ever tip - List contrasts rideshare tipping with traditional taxi and restaurant norms Driver pay mix pre- and post-tip: Hourly wages identical pre-tip and post-tip - Tipping increased labor supply but the wage effect was offset by more empty driving Driver compensation share: About 80% of the take - List explains the typical Uber driver revenue share earlier in the platform’s history Advertising cost per driver on Facebook: About $300 to $500 per driver on average - Lyft acquisition data presented to List Advertising cost per driver on Google: About $700 per driver on average - Lyft acquisition data presented to List Marginal ad cost for Facebook: About $1,000 per driver for the last 25 drivers - The recent Facebook cohort was much more expensive than the average Marginal ad cost for Google: About $750 per driver for the last 25 drivers - The recent Google cohort was cheaper than Facebook at the margin New York ride-hailing supply controls: Implemented in and around New York City - Discussed as a factor affecting prices and supply conditions Taxi medallion history: No change since 1936 (as stated) - Roberts cites fixed taxi medallion supply in New York as contrast to ride-hailing
Pivotal Quotes: "The voltage effect is a description that tells us what happens to a program's effects when we go from the small to the large." — John List: Opening definition of the book’s central concept "If I was primarily interested in horizontal scaling, that works because I think every community can get 30 pretty good teachers. But if I'm interested in vertically scaling within the Chicago market, I think it's going to be very hard for me to get 30,000 really good teachers like those 30." — John List: Explaining why the Chicago Heights pre-K was only half-right for scale "You can add tipping. to the app. Drivers supply more labor hours... And in fact, the amount of driving around with an empty car and its effect on wages exactly offsets the effect of the tips on wages." — John List: Summarizing the Uber tipping experiment and its equilibrium effect
Implications: For firms, governments, and nonprofits, the key lesson is to test for scale from the start, not just success in a pilot. Policies must be designed around realistic inputs, demand response, and marginal costs—or they may fail when expanded.
About EconTalk
EconTalk: Conversations for the Curious is an award-winning weekly podcast hosted by Russ Roberts of Shalem College in Jerusalem and Stanford's Hoover Institution. The eclectic guest list includes authors, doctors, psychologists, historians, philosophers, economists, and more. Learn how the health care system really works, the serenity that comes from humility, the challenge of interpreting data, how potato chips are made, what it's like to run an upscale Manhattan restaurant, what caused the...