EconTalk
EconTalk

Campbell Harvey on Randomness, Skill, and Investment Strategies

Campbell Harvey of Duke University talks with EconTalk host Russ Roberts about his research evaluating various investment and trading strategies and the challenge of measuring their effectiveness. Topics discussed include skill vs. luck, self-deception, the measures of statistical significance, skew

Featured Speakers

Library of Economics and Liberty HostCampbell Harvey Guest

Topics Discussed

Episode Summary

Executive Summary: Russ Roberts and Campbell Harvey discuss how statistical significance can be misleading when many hypotheses are tested. Using physics, a jelly bean cartoon, trading strategies, and sports/medical examples, Harvey argues that data mining inflates false positives, so researchers must adjust thresholds, ask about mechanisms, and consider tail risk and skewness—not just average returns or t-stats.

Main Topics: Statistical significance and the 2-sigma rule (Priority: 5/5): Harvey explains p-values, t-statistics, and the standard 95% confidence / two-sigma convention, while distinguishing statistical from economic significance. Multiple testing, data mining, and false positives (Priority: 5/5): The central theme is that as the number of tests rises, a two-sigma result becomes much less convincing; many 'discoveries' can be flukes. Illustrations from physics and the jelly bean cartoon (Priority: 4/5): The Higgs boson required a five-sigma threshold because trillions of possible false signatures existed; the green jelly bean example shows how repeated testing can manufacture significance. Finance examples: Sharpe ratio and simulated trading strategies (Priority: 5/5): Harvey shows how a seemingly strong trading strategy can be produced by random numbers and still appear significant, illustrating the danger of selecting winners from many trials. Luck versus skill in investing, sports, and prediction (Priority: 4/5): Success streaks in managers, coaches, or forecasters may reflect randomness rather than ability; poor performers may also be victims of bad luck. Tail risk, skewness, and the limits of normality (Priority: 5/5): Harvey criticizes overreliance on the Sharpe ratio and normal-distribution assumptions, arguing that investors must account for asymmetric downside risk and catastrophic losses. Replicability, mechanism, and better research practice (Priority: 5/5): They discuss publication bias, weak replication culture, and the need to ask for causal mechanisms and use stricter thresholds, sensitivity analysis, and better out-of-sample validation.

Key Arguments: A t-statistic above 2 is only meaningful when the researcher has run a single, pre-specified test; with many tests, it can easily be a false positive. Scientific and financial findings should be evaluated in light of how many hypotheses were tried, not just whether one result crossed a conventional threshold. The Higgs boson example shows why fields with huge numbers of possible false positives use much stricter standards like five sigma. A trading strategy can look genuinely impressive, including a 2.5-sigma result, yet be entirely random if it was selected from hundreds of simulated strategies. In the 100,000-email football example, apparent predictive skill is created by repeatedly filtering for the subset that happened to be right by chance. Warren Buffett and elite managers may be skilled, but many top performers in a large population will look exceptional purely because randomness guarantees some winners. The Sharpe ratio is useful but incomplete because it ignores skewness and tail risk; investors should care about catastrophic downside, not just volatility. Big data is valuable, but post hoc storytelling around patterns found by mining data is dangerous unless grounded in first principles and causal logic. Replication and transparency matter: if researchers cannot reveal how many tests they ran or what choices they made, significance claims are less trustworthy.

Data Points: Standard significance threshold: 95% confidence / 2 sigma - Discussed as the conventional cutoff in economics and regression analysis False positive chance under standard rule: 5% - The implied chance that a single significant finding is a fluke Higgs significance threshold: 5 sigma - Used because trillions of collisions/signatures created a vast multiple-testing problem Higgs false-positive probability: about 1 in 1.5 million - Roberts and Harvey contrast this with the 5% level at two sigma Number of collisions/signatures: about 5 trillion - Illustrating the scale of possible false positives in Higgs discovery Number of strategies in paper simulation: 200 - Harvey’s random-number experiment showing that one strategy can look highly significant by chance Simulated strategy significance: 2.5 sigma - The highlighted strategy appears strong but is generated from random returns Football prediction mailings: 100,000 initial emails - Illustration of how repeated filtering can manufacture a winning streak Correct predictors after 10 weeks: 97 people - Those who received 10 consecutive correct picks in the email example Investment bank tests mentioned: 24 different tests - Harvey cites the case where a factor looked significant only after many tries Industrial production lag example: 17th monthly lag - A spurious factor that appeared significant only because many lags were tested Psychology publication rate of positive results: over 92% - Harvey cites this as evidence of publication bias and weak support for null results Manager underperformance example: 2 to 3 years - Russ and Harvey note that skilled managers may be dismissed after a few bad years Example manager performance: 200 to 300 basis points per year - Duke alumnus’s options strategy appeared to outperform by this margin over five years SP 500 reference: benchmark / futures strategy - Used as the comparison for evaluating excess return and manager skill

Pivotal Quotes: "The more tests that you actually do, it's possible to get a result that is just a fluke, something random." — Campbell Harvey: Explaining why multiple testing undermines conventional significance thresholds "You need to take into account that we've got, in this particular situation, 24 different things that have been tried, not one." — Campbell Harvey: Responding to the 17th-lag industrial production example and why the result should not be trusted "I prefer to work on the basis of first principles: what is reasonable." — Campbell Harvey: Describing why post hoc data-mined stories are dangerous even when they fit the numbers

Implications: Listeners should be skeptical of impressive-looking results, especially when many tests or design choices were possible. In finance, medicine, and social science, good practice means stronger thresholds, replication, and causal reasoning—not just p-values.

🔓 Sign Up for Unlimited Episode Search

About EconTalk

EconTalk: Conversations for the Curious is an award-winning weekly podcast hosted by Russ Roberts of Shalem College in Jerusalem and Stanford's Hoover Institution. The eclectic guest list includes authors, doctors, psychologists, historians, philosophers, economists, and more. Learn how the health care system really works, the serenity that comes from humility, the challenge of interpreting data, how potato chips are made, what it's like to run an upscale Manhattan restaurant, what caused the...

View all episodes from EconTalk