Episode Summary
Executive Summary: Stuart Ritchie discusses how science can be distorted by publication bias, p-hacking, hype, weak peer review, and rushed research, especially in psychology, medicine, and COVID-era studies. He argues the scientific method is sound, but academic incentives reward flashy, positive findings over reliable ones, and he outlines reforms like preregistration, registered reports, open data, and better hiring incentives.
Main Topics: Replication crisis and the psychic-powers anecdote (Priority: 5/5): Ritchie opens with his replication of Darrell Bem’s controversial psychic-powers study, illustrating how a striking result in a prestigious journal failed to replicate and was rejected when submitted as a replication. Systemic problems across disciplines (Priority: 5/5): The conversation broadens from psychology to economics, computer science, medicine, physics, and chemistry, arguing that irreproducible results, fraud, and weak methods are widespread rather than isolated to one field. How science is supposed to work vs. how it often works (Priority: 4/5): Whipple and Ritchie contrast the ideal scientific pipeline—planned analysis, statistical testing, peer review—with real-world practices where analysis is often ad hoc and peer review is overloaded and inconsistent. P-hacking and publication bias (Priority: 5/5): Ritchie explains how researchers can unconsciously or deliberately run many analyses until a significant result appears, and how journals’ preference for positive findings creates a literature skewed toward false positives. Hype, spin, and communication distortions (Priority: 4/5): Even when data are legitimate, authors and press releases can overstate findings, infer causality from correlation, or generalize from small or non-human samples, pushing results beyond what the evidence supports. Reforms and repair mechanisms (Priority: 4/5): Ritchie highlights preregistration, registered reports, data sharing, and structural incentives that reward transparency and robustness rather than paper counts and glamour journals. COVID-19 as a stress test for science (Priority: 4/5): The pandemic accelerated publication speed and exposed weaknesses in review and verification, including high-profile retractions and a flood of low-value or unreliable research.
Key Arguments: Prestigious journals can publish extraordinary claims while refusing to publish replications, which incentivizes novelty over truth. Psychology has been studied most systematically, but similar reproducibility and publication-bias problems exist in economics, computer science, and medicine. Scientists often analyze data ad hoc instead of following a preplanned analysis, making it easier to find spurious significance. P-hacking inflates the rate of positive findings because repeated testing on the same dataset can produce chance results that look meaningful. Publication bias means null or negative studies disappear at every stage: publication, writing, and citation. Scientific hype is often generated by the scientists themselves through paper text and press releases, not only by journalists or university PR staff. The scientific method remains valid; the problem is the incentive structure of academia, journals, and funding systems. Preregistration and registered reports can reduce bias by locking in analysis plans before data collection. Open data and code sharing improve reproducibility and make errors easier to catch. COVID-era speed increased the volume of research but also increased mistakes, retractions, and low-value papers.
Data Points: Replication study outcome: No positive result - Ritchie’s replication of Bem’s psychic-powers experiment found that participants did not remember future words. Original journal policy: 0 replication studies accepted - The Journal of Personality and Social Psychology reportedly would not consider replication studies under any circumstances. Positive findings in psychology: 90-something percent - Ritchie says psychology papers are overwhelmingly positive, suggesting publication bias. Clinical trial evidence base: About 50-50 positive vs negative - In the antidepressant trial example, registered studies were roughly evenly split before publication bias altered the record. Published negative studies after filtering: About 5 - Ritchie describes a study where many negative trials largely disappeared through publication and spin. Total registered antidepressant studies: 100 - The example referenced a study of roughly one hundred registered antidepressant trials. Peer review example quote count: 2 quoted reviews - Ritchie and Whipple cite especially harsh peer-review comments to illustrate the system’s tone and limits. COVID-related retractions: 2 top-journal papers retracted - Papers in The Lancet and NEJM on hydroxychloroquine were retracted after data concerns.
Pivotal Quotes: "We do not accept replication studies, we do not consider publication of replication studies under any circumstances." — Tom Whipple quoting the journal’s response: The rejection of Ritchie’s replication of the psychic-powers paper. "I really like it when my students just gather a whole bunch of data with a whole bunch of stuff, and then they just keep analyzing it until they find something, and then we publish that." — Stuart Ritchie: Used to explain p-hacking and Brian Wansink’s problematic research culture. "The incentives are all wrong." — Stuart Ritchie: His core diagnosis of why academia rewards publication quantity and glamour over reliability.
Implications: Listeners should be more skeptical of flashy studies, especially from high-pressure fields or rapid news cycles. The broader research system needs preregistration, transparency, and better incentives if science is to become more reliable and reproducible.