Episode Summary
Executive Summary: The episode examines the so-called replication crisis in science, arguing that failed replications often reveal nuance rather than fraud. Using psychology examples—especially the Asian women/math stereotype study—it shows how different labs can get different results due to context, design, and randomness, and why science advances by accumulating evidence rather than delivering final truth.
Main Topics: The replication crisis in science (Priority: 5/5): The episode introduces widespread concern that many published findings, especially in psychology and medicine, do not hold up when independently repeated. Examples of controversial and flawed studies (Priority: 4/5): Cases like broken sidewalks and racism, gay contact reducing anti-gay bias, and the bilingual advantage are used to show how some celebrated findings were later undermined by fraud, retraction, or incomplete reporting. Why reproducibility matters (Priority: 5/5): Brian Nosek explains that a scientific claim becomes credible when other researchers can reproduce it using the same protocol and procedures. The Asian women and math stereotype study (Priority: 5/5): Todd Patinsky and Margaret Shee’s study found that reminding Asian American women of gender hurt math performance, while reminding them of Asian identity improved it; it became a famous stereotype-threat/stereotype-boost example. Competing replications and contextual differences (Priority: 5/5): Replications at Georgia Southern and UC Berkeley produced different outcomes, raising questions about location, experimenter gender, and whether the studies were truly identical. What failed replication really means (Priority: 5/5): Experts argue that replication is not primarily a fraud detector; different results can reflect sample size, randomness, moderators, or boundary conditions rather than invalid science. Science as cumulative, probabilistic knowledge (Priority: 4/5): The episode concludes that science is about building more precise estimates over time, not finding permanent final answers.
Key Arguments: Many high-profile findings in psychology and adjacent fields have not replicated, which has prompted concern about the reliability of published science. Some failures stem from outright fraud or retracted papers, but others reflect incomplete reporting, selective publication, or differences in context rather than deception. A study becomes more credible when independent researchers can reproduce it with similar methods and similar results. Exact replication is often unrealistic in social science because small procedural changes, location, population, and experimenter characteristics can alter outcomes. Different replication results do not automatically mean the original study was false; they may indicate effect-size overestimation, randomness, or important moderating variables. Replication should be used to refine understanding and estimate the size and conditions of effects, not only to expose misconduct. Science is incremental: each study is one data point, and truth emerges from many dots over time rather than a single decisive experiment.
Data Points: Studies replicated in OSC project: 100 - Psychology experiments the Center for Open Science recruited teams around the world to replicate Replication success rate: less than half - Brian Nosek reported reproducing original results in fewer than half of cases across five criteria Year of replication report: 2015 - Nosek’s large replication effort was published that year Broken sidewalks study year: 2011 - Dutch researchers published the claim that broken sidewalks encourage racism Gay marriage contact study year: 2009 - A Science paper claimed that personal contact with a gay person could change anti-gay-marriage views Bilingual advantage experiments: 4 - The paper had four experiments, but only one showed a benefit and only that one was published Experiments showing bilingual advantage: 1 - Only one of four experiments supported the claim Harvard replication sites: 2 - The Asian women/math study was replicated at Georgia Southern University and UC Berkeley Original study publication year: 1999 - The stereotype/Asian women math finding became a textbook example after publication SAT prep example effect sizes: 15% vs 3% - Eric Bradlow used this to explain how different studies may vary yet still support a real effect Mice recovery times: 35 minutes vs 16 minutes - NIH example showing how cage wood shavings affected sedation recovery time
Pivotal Quotes: "One of the key factors of science is that a claim becomes a credible claim by being reproducible." — Brian Nosek: Explaining why replication is central to scientific credibility "Science is the slow march of accumulating evidence." — Brian Nosek: Describing replication as an ongoing process of refinement rather than a final verdict "If you want answers that never change, definitive conclusions, and final truths, odds are you don't want to ask a scientist." — Shankar Vedantam: Concluding the episode’s view of science as probabilistic and evolving
Implications: Listeners should treat scientific findings as provisional and context-dependent. For research and publishing, the episode supports preregistration, more replications, and better incentives for careful, cumulative science.
About Hidden Brain
Why do I feel stuck? How can I become more creative? What can I do to improve my relationships? If you’ve ever asked yourself these questions, you’re not alone. On Hidden Brain, we help you understand your own mind — and the minds of the people around you. (We're routinely rated the #1 science podcast in the United States.) Hosted by veteran science journalist Shankar Vedantam.