Episode Summary
Executive Summary: The episode examines the replication crisis in social science through DARPA’s SCORE project and a separate replication effort in economics and political science. The guests argue that replication is improving, but only because the field is increasingly emphasizing data/code sharing, better training, journal and funder policies, and emerging AI tools that can both complicate and strengthen reproducibility.
Main Topics: The replication crisis in social science (Priority: 5/5): The discussion frames replication as the core test of scientific trustworthiness and explains why many published findings cannot be assumed true without follow-up confirmation. DARPA’s SCORE project (Priority: 5/5): The guests describe SCORE as a large-scale, cross-disciplinary replication effort intended not only to re-test findings but also to create ground truth for AI tools that assess confidence in research. Barriers to replication: data, code, and methods (Priority: 5/5): A major theme is that many papers lack sufficient sharing of data, methods, and analysis code, making reproduction difficult and limiting confidence in results. Evidence that research practices are improving (Priority: 4/5): Brodur argues that more recent economics and political science papers show better data sharing and higher robustness than earlier work, though mistakes and non-replicable results remain common. Human error and research incentives (Priority: 4/5): The guests stress that most replication failures are not necessarily fraud; they often stem from honest mistakes, rushed workflows, and incentives that reward positive results over transparency. AI’s double role in replication (Priority: 4/5): AI may worsen uncertainty by enabling machine-generated research text, but it may also help automate reproduction, improve documentation, and test multiple plausible analytical strategies. What listeners should trust (Priority: 3/5): Brodur recommends skepticism toward single studies and greater confidence only after results are independently repeated multiple times, reinforcing science as a cumulative process.
Key Arguments: Replication is essential because publication alone does not establish truth; research findings need follow-up confirmation to build trust. SCORE was unusually broad, examining 10 years of journals across 62 journals and using replication as a basis for developing AI confidence-assessment tools. A lack of shared data, code, and methodology is one of the biggest obstacles to verifying results and is not yet the norm in many fields. Replication failures can arise from many non-fraud reasons, including missing data, coding errors, weak robustness, and differences in new datasets. Recent work in economics and political science appears stronger than older work, with improved data sharing and higher robustness rates. Training and institutional norms matter: researchers, journals, funders, and universities all influence whether rigorous reproducibility practices become standard. AI could automate parts of replication and robustness checking, but it could also make it harder to tell whether a paper was written or analyzed by a human. Consumers of research should wait for repeated independent confirmation before treating a striking new result as reliable.
Data Points: SCORE journal coverage: 62 journals - Tim Arrington described SCORE as analyzing papers across 62 social-behavioral science journals. Time span analyzed: 10 years - The SCORE project looked at a decade of journal publications. Replication rate in SCORE: about 50% - Arrington said researchers could only replicate half the papers analyzed in the SCORE project. Recent paper window: 2022–2023 - Abel Brodur said his newer replication study focused on papers published in 2022 and the end of 2023. Coding errors: 15–20% of papers - Brodur reported finding coding errors in a notable minority of recent economics/political science papers. Robust results: about 75% - Brodur said results were robust roughly three-quarters of the time in his newer replication work.
Pivotal Quotes: "It got published. It must be true. But that's not how science works." — Dr. Tim Arrington: He was explaining why publication status cannot be treated as proof of validity when data and methods are not fully shared. "If I see a new result, something like innovative, like the first time that I hear about something, I don't believe it." — Dr. Abel Brodeur: He described his personal rule for trusting findings only after repeated independent replications. "Science is a process." — Dr. Tim Arrington: He emphasized that each paper is only one step in an evolving effort to approach truth.
Implications: Replication standards are tightening, but trust still depends on sharing data/code, better training, and stronger institutional incentives. Listeners should treat exciting new findings as provisional until they are independently repeated.