EconTalk
EconTalk

Brian Nosek on the Reproducibility Project

Brian Nosek of the University of Virginia and the Center for Open Science talks with EconTalk host Russ Roberts about the Reproducibility Project--an effort to reproduce the findings of 100 articles in three top psychology journals. Nosek talks about the findings and the implications for academic pu

Featured Speakers

Library of Economics and Liberty HostRuss Roberts GuestBrian Nosek Guest

Topics Discussed

Episode Summary

Executive Summary: Russ Roberts and Brian Nosek discuss the Reproducibility Project in psychology: what reproducibility means, why published findings can diverge from truth, and what a large-scale replication effort found. The project showed substantial decline from original to replication effects, exposing publication bias, low power, and flexibility in analysis, while also highlighting field differences and the need for open, transparent science.

Main Topics: What reproducibility means in science (Priority: 5/5): Nosek distinguishes between rerunning analyses on the same data and true replication using new data under comparable conditions. Why published results may not be reliable (Priority: 5/5): The conversation emphasizes incentives for novelty, prestige, and statistically significant findings, which can distort the published literature away from truth. The Reproducibility Project: design and scope (Priority: 5/5): The project sampled 2008 studies from three top psychology journals, used original materials when possible, and coordinated large-scale replications with transparent procedures. What the replication results showed (Priority: 5/5): Replication effects were smaller than originals, with far fewer significant results; multiple criteria were used to judge success because replication is not a single binary concept. Interpreting decline effects and false positives (Priority: 4/5): They discuss publication bias, underpowered studies, exploratory data mining, and how these can produce inflated original effects that later shrink on replication. Field differences and future open science efforts (Priority: 4/5): Cognitive psychology appeared to replicate better than social psychology, and Nosek describes ongoing replication projects in cancer biology and other fields.

Key Arguments: Reproducibility is a core norm of science because credibility should come from independent confirmation, not authority or prestige. The published literature is likely biased upward because researchers are rewarded for positive, novel, statistically significant results, not for accuracy. Exploratory analysis is valuable, but conclusions from exploratory work must be tested on new data rather than treated as confirmatory. The project intentionally used transparency, original-author input, and open materials to reduce hostility and improve the fairness of replications. A replication outcome is not always a simple yes/no; effect size, statistical significance, and subjective judgment can each tell a different story. Smaller replication effects may reflect publication bias, but could also reflect context sensitivity, especially in social psychology. The project should be viewed as a large empirical probe into reproducibility, not a final verdict on an entire field. Open science infrastructure and larger cross-disciplinary replication efforts are needed to better estimate where reproducibility problems are strongest.

Data Points: Year sampled: 2008 - The project selected studies from three psychology journals published in 2008. Journals sampled: 3 - Psychological Science, JPSP, and JEP: Learning, Memory, and Cognition. Completed replications: 100 - The final Science report included 100 completed replication studies. Original studies with significant results: 97% - Nearly all original studies in the sample reported statistically significant findings. Replication studies with significant results: 36% - Only about a third of replications achieved statistical significance. Subjective replication success: 39% - Replication teams judged 39% of effects as reproducing the original result. Effect size of replications: About half of originals - Average replication effects were described as half the magnitude of original effects. Original effect sizes inside replication CI: 47% - Less than half of original effects fell within the replication effect's 95% confidence interval. Combined result under no-bias assumption: 68% significant - When original and replication results were combined, 68% were significant if no bias in original results is assumed. Scope of eligible studies: About 165 eligible out of 488 possible - Nosek described a constrained sampling frame from the 2008 journal issues. Started replications: 113 - More studies were started than finished. Cognitive vs. social psychology: About 2x higher replication rate in cognitive psychology - Cognitive findings replicated at roughly twice the rate of social psychology findings.

Pivotal Quotes: "published and true are not synonyms" — Russ Roberts: Roberts emphasizes that journal publication does not guarantee truth. "Reproducibility is central to science" — Brian Nosek: Nosek explains why independent confirmation is foundational to scientific credibility. "the only way you can get statistical significance is to take advantage of chance" — Brian Nosek: He describes how underpowered studies and publication bias can inflate original findings.

Implications: The episode argues for preregistration, transparency, and more replications across disciplines. For researchers and readers, it is a warning that published findings can overstate reality and that stronger evidence practices are needed.

From the Transcript

I try to get published. And if that isn't a complete representation of how I got to those findings, then what's in the published literature could be more beautiful than what reality is. Aaron Powell, Jr.: So, in the paper you wrote a while back that we talked about three years ago, there are two things I just want to mention again because they're so important. One is the line: published and true are not synonyms, which is hard for people to accept. I think John. Journalists certainly have an incentive to ignore that truth. So they publish lots of things that are published results in peer-reviewed journals because they're really interesting and people want to read about them. Whether they're true or not, it's a different question. But just quickly retell the research finding you had about shades of gray that was in that original paper. Yeah. So Matt and Motel were interested in a very fascinating area of research in psychology right now.

Russ Roberts · at 7:46

I'm going to say this in a negative way on purpose. Why would you waste your time replicating all these results? We've already found them. What are some of the issues? We talked about this three years ago, but I want to review them. Why would you be suspicious or concerned about some of the findings in the finest peer-reviewed journals in your field or mine? Well, the general answer is that reproducibility is central to science, right? So a scientific claim doesn't gain credibility because Of an authority saying this is true, you have to believe it, or because that person has a good reputation. Scientific claims gain credibility by the ability for them to be independently reproduced. Someone else can follow the same procedure and find the same result. So the credibility is within the evidence itself, not within the generator of needs. So, that is a core principle of how scientific claims become credible, means that having reproducible research is.

Brian Nosek · at 3:52

In the kinds of research that we do. This is a pervasive problem that's been discussed since the 1960s. But the consequence of those two things happening simultaneously, low-powered research and requirement for statistical significance, means that the only way you can get statistical significance is to take advantage of chance and happen to observe larger than reality, larger than what are real effects. That just because I run five studies, I am investigating. A true effect, one of those will happen to be larger than it really is and obtain that statistical significance, and that's the one that gets published. So, if that is occurring at a pervasive scale, then the results of the reproducibility project are exactly what you would expect as a consequence, which is most research is actually when you just do it and report it, regardless of whether it's significant or not, is going to estimate smaller effects than those few that get through the publication library.

Brian Nosek · at 45:38
🔓 Sign Up for Unlimited Episode Search

About EconTalk

EconTalk: Conversations for the Curious is an award-winning weekly podcast hosted by Russ Roberts of Shalem College in Jerusalem and Stanford's Hoover Institution. The eclectic guest list includes authors, doctors, psychologists, historians, philosophers, economists, and more. Learn how the health care system really works, the serenity that comes from humility, the challenge of interpreting data, how potato chips are made, what it's like to run an upscale Manhattan restaurant, what caused the...

View all episodes from EconTalk