Episode Summary
Executive Summary: Susan Athey argues that modern econometrics is converging with machine learning: use data-driven methods to improve prediction and control selection, but preserve causal inference through careful design, sample splitting, and robustness checks. The conversation contrasts policy evaluation’s identification problems with tech firms’ rigorous A/B testing, highlighting both the promise and limits of experiments, big data, and algorithmic model selection.
Main Topics: Why policy effects are hard to measure (Priority: 5/5): Athey explains that comparing places or times directly confounds causation with correlation because jurisdictions adopting policies like the minimum wage differ systematically from those that do not. Difference-in-differences, placebo tests, and synthetic controls (Priority: 5/5): The discussion covers standard causal inference tools used to estimate counterfactuals, including comparing pre-trends, running placebo tests, and building weighted synthetic control groups from multiple comparison units. Machine learning as high-dimensional control selection (Priority: 5/5): Athey argues machine learning is useful for choosing among many possible covariates and capturing complex, multidimensional patterns that traditional regressions handle poorly. Data mining, p-hacking, and overfitting (Priority: 5/5): Russ Roberts and Athey discuss the danger of searching many specifications until one looks significant, which invalidates p-values and can produce false findings if not handled transparently. Sample splitting and honest inference (Priority: 4/5): A central proposed fix is to use one sample to select variables/models and another to estimate effects, or repeated-split variants that preserve valid inference while still using machine learning. A/B testing in tech firms as rigorous experimentation (Priority: 4/5): Athey describes how Google, Amazon, Bing, and Facebook use randomized trials to evaluate product changes, enabling strong causal claims within their own platforms. Limits of experiments and the short-term bias of firms (Priority: 4/5): Even highly scientific tech firms may optimize for short-run outcomes because long-run effects are costly and hard to measure, leading to systematic biases in innovation.
Key Arguments: Comparing high- and low-minimum-wage cities is not enough because policy adoption is endogenous; the places that raise wages often differ in growth, rent levels, labor scarcity, and customer price sensitivity. Policy evaluation requires a counterfactual: what would have happened without the intervention, which is why quasi-experimental designs are necessary. Difference-in-differences works by using untreated areas or pre-policy trends as a control for the treated area’s underlying trend. Synthetic controls improve credibility by combining multiple comparison units weighted to match the treated unit on many dimensions and pre-trends. Machine learning is valuable for selecting from large sets of covariates, especially when the number of potential controls exceeds the number of observations. Off-the-shelf machine learning can reproduce classic econometric errors if it is used without attention to causal identification and inference. Data mining becomes misleading when researchers try many specifications and report only the best one; the reported p-values no longer mean what they claim. Sample splitting makes model selection and estimation separate, reducing overfitting and restoring valid standard errors and p-values. A/B tests provide near-perfect causal identification inside firms because random assignment separates correlation from causation. Tech-company experiments are often highly rigorous but mostly answer narrow, context-specific questions rather than general scientific laws. Economists and machine-learning researchers should collaborate: econometrics contributes causal logic and inference; machine learning contributes flexible prediction and large-scale model selection. Both academia and industry systematically make different mistakes: academics may overstate fragile causal claims, while tech firms may overemphasize short-term measurable gains and miss long-term innovation.
Data Points: Date of interview: June 18, 2016 - Opening of the EconTalk episode Minimum wage example: $15, $10, $8 - Roberts and Athey discuss comparing cities/states with different wage floors Historical minimum wage study: Card and Krueger: Pennsylvania vs. New Jersey - Cited as a landmark difference-in-differences-style analysis Model selection scale: 10,000 models - Athey suggests letting a machine explore far more specifications than a human researcher typically would Hypothetical drug-trial sample: 1,000 patients and 1,000 characteristics - Used to show how data mining can always find spurious subgroups Strong outcomes subgroup: 15 patients - Illustrates how a small lucky subgroup can falsely appear to identify treatment effects Tech experimentation scale: 10,000 or more randomized controlled trials per year - Athey’s description of Google’s experimentation culture Short-term vs realized impact: 300% predicted improvement vs 10% actual change - Used to illustrate how summing short-term experiment gains can overstate long-run business results Platform experiment example: 100s to 10,000s of users / many repeated exposures - Referenced in discussing correlated observations in A/B testing and the need for valid statistics
Pivotal Quotes: "The big fallacy there is that the places that have chosen to raise the minimum wage are not the same as the places that have not." — Susan Athey: Explaining why simple cross-sectional comparisons of minimum wage levels and employment are misleading "If you look hard enough, you'll find the result you're looking for." — Susan Athey: Describing the danger of data mining and p-hacking when many specifications are tried "Machine learning doesn't solve any of those problems." — Susan Athey: Stressing that better prediction tools do not by themselves resolve causal identification issues
Implications: For policy, better data and machine learning can improve causal inference if used honestly. For firms, A/B testing is powerful but can bias toward short-term wins. For researchers, the future is hybrid: machine-learning-driven selection plus econometric identification and robust inference.
About EconTalk
EconTalk: Conversations for the Curious is an award-winning weekly podcast hosted by Russ Roberts of Shalem College in Jerusalem and Stanford's Hoover Institution. The eclectic guest list includes authors, doctors, psychologists, historians, philosophers, economists, and more. Learn how the health care system really works, the serenity that comes from humility, the challenge of interpreting data, how potato chips are made, what it's like to run an upscale Manhattan restaurant, what caused the...