Episode Summary
Executive Summary: The episode explores how machine learning is reshaping drug discovery, with Daphne Koller explaining why her company InSitro focuses on integrating computational models with wet-lab biology rather than “AI” hype. The conversation covers data scale in biology, causal inference using human genetics and genome editing, disease selection strategy, industry skepticism, and the need for interdisciplinary teams to turn cell-level insights into better therapies.
Main Topics: Why machine learning is changing drug discovery (Priority: 5/5): Russ Altman frames drug discovery as a longstanding challenge now increasingly influenced by computation, data science, and AI/ML methods that help find targets and prioritize interventions. AI vs. machine learning (Priority: 5/5): Koller distinguishes machine learning from AI, arguing that drug discovery is better described as applying learning methods to large data rather than trying to mimic human intelligence. Why biology is becoming data-rich enough for ML (Priority: 5/5): Koller explains that single-cell technologies, imaging, and large biobanks have made biological and patient data much larger and more suitable for high-end machine learning. How InSitro finds drug targets (Priority: 5/5): The company combines in silico and in vitro methods to model disease in cells, detect phenotypic differences, and identify interventions that revert unhealthy cell states toward normal. Causality, genetics, and genome editing (Priority: 4/5): The discussion addresses the danger of confusing correlation with causation and highlights human genetics and CRISPR-based perturbations as key tools for causal inference. Choosing disease areas with tractable cell models (Priority: 4/5): InSitro’s initial focus on liver disease and brain disorders reflects the need for scalable cellular systems and strong genetic components, while more complex multi-organ diseases are harder to model. Building a hybrid biotech-technology culture (Priority: 4/5): Koller discusses the challenge and opportunity of blending software, ML, stem-cell biology, and translational science into one team with a shared vocabulary and mutual respect.
Key Arguments: Machine learning is more appropriate than the broader term AI for this work because the goal is to solve hard biomedical problems using data, not replicate human intelligence. Biology has reached a scale where ML can be useful, especially through single-cell measurements and large population biobanks. Unbiased, high-dimensional cell profiling can reveal disease patterns that humans may miss because people tend to focus on only a few variables with strong priors. Human genetics offers a causal bridge between molecular signals and disease outcomes, making it valuable for target selection. Genome editing allows controlled perturbations in cells, helping determine whether observed changes are actually causal. Drug discovery remains extremely difficult: even improved methods still operate in a field with very high failure rates and long timelines. Success in this space requires integrating computational and experimental expertise rather than treating data science and biology as separate functions. A healthy culture in biotech-ML organizations depends on shared respect, cross-training, and employees who can bridge the wet-lab/dry-lab divide.
Data Points: Drug discovery success rate: About 5% - Koller cites the approximate end-to-end success rate for drug development programs. Drug discovery failure rate: About 95% - Used to emphasize how difficult it is to bring a drug from concept to approval. Timeline for AI winters: Two or three AI winters - Koller references her experience through multiple periods of disappointment in AI hype cycles. Scale of early biology datasets: A couple hundred samples - Describes the small size of many early biological machine learning datasets. Population biobank scale: Hundreds of thousands of patients - Refers to resources such as UK Biobank and All of Us as enabling larger-scale ML in human health. Single-cell scale: Millions to hundreds of millions of cells - Highlights how single-cell technologies create sufficiently large datasets for modern machine learning. Company founding timeline: Around 2016 - Koller says her move into founding InSitro began around this time. Company age at discussion: Two and a half years - She notes that industry skepticism was stronger when the company began roughly this long ago. Interdisciplinary program age: Almost 20 years - Russ and Daphne reference the biomedical computation major at Stanford that they helped establish.
Pivotal Quotes: "I don't think of what we're doing as AI for drug discovery, but rather machine learning for drug discovery." — Daphne Kohler: Koller opens by distinguishing the kind of computational work InSitro does from broader AI claims. "Let the computer figure it out." — Daphne Kohler: She explains the value of unbiased, high-dimensional measurement in discovering disease-relevant cellular patterns. "Biology is really hard." — Daphne Kohler: Koller underscores the complexity of the field and cautions against overpromising on ML’s impact.
Implications: Machine learning can materially improve target discovery and biological insight, but only when paired with rigorous biology, causal validation, and realistic expectations. The future of drug discovery likely depends on hybrid teams and large-scale cellular and human datasets.
About The Future of Everything
Host Russ Altman, a professor of bioengineering, genetics, and medicine at Stanford, is your guide to the latest science and engineering breakthroughs. Join Russ and his guests as they explore cutting-edge advances that are shaping the future of everything from AI to health and renewable energy. Along the way, “The Future of Everything” delves into ethical implications to give listeners a well-rounded understanding of how new technologies and discoveries will impact society. Whether you’re a ...