Super Data Science: ML & AI Podcast with Jon Krohn
Super Data Science: ML & AI Podcast with Jon Krohn

703: How Data Happened: A History, with Columbia Prof. Chris Wiggins

Statistics history, interdisciplinarity, and data and society. Chris Wiggins talks with Jon Krohn about the power dynamics of data, the transformation of the field of biology through data-driven approaches to genetic sequencing, and the New York Times’ data science team’s cutting-edge approach to ac

Featured Speakers

Jon Krohn HostChris Wiggins Guest

Topics Discussed

Episode Summary

Executive Summary: Dr. Chris Wiggins traced data science from its 18th-century statistical roots to today’s AI era, arguing that the field is inseparable from history, ethics, and power. He made the case for humanities training in data curricula, explained why Bayes and statistics became controversial, and described how data reshapes society, journalism, biology, and corporate decision-making.

Main Topics: History of data and statistics (Priority: 5/5): Wiggins explains how statistics emerged from statecraft and the scientific revolution, then evolved into a tool for describing and prescribing social policy, not just counting. Why humanities belong in data science (Priority: 5/5): He argues that data scientists need humanities training for communication, strategy, historical context, and ethical reasoning, since technical skill alone is insufficient. Data science as power rearrangement (Priority: 5/5): Borrowing from cryptography, Wiggins says data science changes who can do what, making it inherently political in the sense of power dynamics and institutional influence. Three eras of data in How Data Happened (Priority: 4/5): He outlines the book’s structure: statistics as an academic field, WWII/Bletchley Park and computing, and the modern battle over data ethics and corporate power. Bayes, controversy, and statistical philosophy (Priority: 4/5): He explains why Bayesian methods were historically rejected by dominant statisticians and why their use still raises questions about priors, subjectivity, and truth-making. Data ethics and applied governance (Priority: 5/5): He frames ethics as a practical organizational challenge, especially for corporations, and recommends principles drawn from human-subjects research for responsible deployment. Computational biology and modern data infrastructure at The New York Times (Priority: 4/5): Wiggins discusses how sequencing and machine learning transformed biology, and how the Times’ data stack shifted from Hadoop/MapReduce toward BigQuery, Python, Airflow, and containers.

Key Arguments: Data scientists should study history because understanding how a field emerged is a form of root-cause analysis for present-day problems. Humanities training matters because communication, ethical judgment, and strategy are central to data science work, not optional extras. Data science rearranges power by changing who can do what, which makes it political in the sense of power relations, not party politics. Data and algorithms are not neutral or raw; they are shaped by subjective design choices about collection, storage, and modeling. Bayesian inference is mathematically ordinary but was historically controversial because it requires explicit prior belief. Modern data ethics needs more than a checklist; it needs organizational processes, audits, and accountable decision points. Biology became genuinely data-driven after large-scale sequencing made statistical and machine-learning approaches indispensable. The New York Times’ move to managed cloud infrastructure improved reliability and speed compared with earlier MapReduce/Hadoop workflows.

Data Points: Year the word "statistics" entered English: 1770 - Used by Wiggins as a natural starting point for the history of data and statistics. Year of the first freely living organism sequenced: 1995 - He cites Haemophilus influenzae as the turning point in biology becoming data-driven. Duration of his tenure at The New York Times: nearly a decade - He notes he has been chief data scientist there for almost ten years. Year the NYT digital paywall launched: 2011 - He references the shift from ad-based revenue toward subscriptions. Year of Bayes’ original essay: posthumous publication after Bayes died - He explains that Bayes wrote an essay on miracles that was later published by a friend. Audience-supported tech stack at The New York Times: Python, scikit-learn, Go, BigQuery, Airflow, Docker/containerization - He describes the modern stack used by the NYT data science team. Bletchley Park intelligence impact window: World War II era - He frames the birth of programmable computing as a data problem during WWII. Number of books published in under a year: 2 - How Data Happened and Data Science in Context were both released within months.

Pivotal Quotes: "Data science rearranges power, it's changing who can do what from what." — Chris Wiggins: Explaining why data science is politically consequential in the power-dynamics sense. "The data are never raw." — Chris Wiggins: Discussing the subjective choices involved in data collection, curation, and analysis. "We should be able to understand this craft using numbers and later using statistics." — Chris Wiggins: Describing the historical shift toward quantifying domains that were once understood qualitatively.

Implications: Listeners should see data science as a historical, social, and ethical discipline—not just a technical one. Future practitioners need communication, judgment, and governance skills to build trustworthy systems and resist harmful uses of power.

🔓 Sign Up for Unlimited Episode Search

About Super Data Science: ML & AI Podcast with Jon Krohn

View all episodes from Super Data Science: ML & AI Podcast with Jon Krohn