The a16z Podcast
The a16z Podcast

a16z Podcast: Deep Learning for the Life Sciences

with Vijay Pande (@vijaypande) and Bharath Ramsundar Deep learning has arrived in the life sciences: every week, it seems, a new published study comes out... with code on top. In this episode, a16z General Partner Vijay Pande and Bharath Ramsundar t...

Featured Speakers

a16z HostBart Ramsundar GuestVijay Pandey Guest

Topics Discussed

Episode Summary

Executive Summary: The episode argues that deep learning is finally becoming useful in life sciences because better hardware, cloud software, open-source code, and new biological representations let models learn directly from raw data. Vijay Pandey and Bart Ramsundar explain how this is democratizing biology, shifting it toward an open, hacker-like ecosystem, while also noting major gaps in standards, causality, reproducibility, and biological tooling.

Main Topics: Why deep learning matters in biology now (Priority: 5/5): The speakers explain that machine learning can learn from raw biological data instead of relying only on hand-built theoretical models, making this the right moment for breakthroughs in genomics, chemistry, microscopy, and drug discovery. Representation learning as the core breakthrough (Priority: 5/5): A major theme is that success depends on choosing or learning the right representation of molecules, proteins, or biological signals; graph-based methods and learned embeddings are presented as more powerful than fixed human-designed encodings. Open source, democratization, and the hacker ethos (Priority: 4/5): The conversation emphasizes that GitHub, Arxiv, and accessible code have lowered barriers to entry, enabling self-taught practitioners and broad participation in biology much like open-source software transformed computing. Limits of current biological ML (Priority: 5/5): The speakers stress that biology has not yet reached the same maturity as image or text ML because the field still lacks foundational building blocks, robust data infrastructure, and standardized ways to operationalize models. Causality, interpretability, and trust (Priority: 5/5): They discuss the need to distinguish correlation from causation, use time-series and real-world evidence, and build interpretability tools so scientists can trust model outputs rather than treat deep learning as a black box. Organizational and cultural barriers (Priority: 4/5): Existing biotech and pharma organizations often struggle to adopt AI because of legacy systems, siloed teams, and risk-averse cultures; the speakers argue that change will come from new companies and a new generation of hybrid-trained talent. Toward open-source biology and decentralized research (Priority: 4/5): The episode ends by imagining decentralized communities, simulation infrastructure, and even market-like mechanisms for evaluating biological hypotheses, pointing toward a future where biology becomes programmable and more open.

Key Arguments: Machine learning is valuable in biology because it can learn directly from raw data, unlike traditional theory-heavy approaches that often fail to capture biological complexity. The decisive shift is not just algorithms, but the combination of GPUs, cloud computing, software stacks, and open-source knowledge that makes deep learning practical and teachable. Biology depends on representation learning: the right mathematical form can make a hard problem tractable, just as converting Roman numerals into decimal simplifies arithmetic. Graphs are a promising representation for molecules because they preserve structure better than naive encodings, enabling graph neural network approaches. Biology is still behind images and text in ML maturity by roughly several years, so the field is developing its own methods rather than merely copying computer vision. Democratization in ML is partly an illusion because core model families already exist; biology still lacks equivalent “Lego pieces” that newcomers can reliably assemble. Open-source ecosystems in software show how companies can gain influence by releasing tools, and similar dynamics may push bio companies toward openness. Pharma and biotech adoption is slowed by legacy infrastructure, toolchain incompatibilities, and organizational silos between computational and scientific teams. The highest-value biological ML work requires combining world-class biological intuition with machine learning skill; neither discipline alone is enough. Interpretability matters because models can latch onto spurious shortcuts, such as scanner labels or rulers in images, rather than the biology itself. Causality is increasingly accessible through time-series data, real-world evidence, probabilistic models, and simulations, especially in healthcare and drug discovery. A future bio-AI stack may resemble self-driving-car simulation loops, where models test hypotheses, receive feedback, and improve continuously. Open-source biology could emerge as a decentralized brain trust of contributors, but reproducibility and trust mechanisms remain major unresolved challenges. A possible near-term application is pricing or de-risking biological assets so that investments can be made more granularly, allowing more hypotheses to be explored. The book aims to give practitioners a practical lens for deciding when a problem is actually suitable for ML and when human biological expertise is needed.

Data Points: Publication cadence: paper a week - Describes how fast impactful machine-learning research in life sciences is now appearing compared with earlier slower cycles. Earlier publication cadence: paper a quarter or a paper a year - Used to contrast current speed with the past pace of breakthroughs in the field. Time gap in biological ML maturity: about 5 years behind - Speaker compares biology’s deep learning progress to images and text, which are said to be several years ahead. Open-source software cost: $500 - Referenced as a historical example of how academic software used to be sold, illustrating improved access today. Training horizon for industry change: 10 years - Estimated time needed for a new generation of students and practitioners to shift biotech culture toward AI-native methods. Best biologist market rate comparison: lower than an introductory/intermediate front-end developer - An anecdote about mispriced talent in biotech, highlighting labor market imbalance. Server scale analogy: 10,000 servers - Used to illustrate how cloud computing can parallel many human workers for machine learning tasks. Example of simple human task: in a second - Rule of thumb offered: if a human can do a perceptual task in about a second, deep learning may be able to do it too. Model performance concern: close to 1.0 - Refers to suspiciously high AUC/accuracy that may indicate a spurious shortcut rather than real biological understanding.

Pivotal Quotes: "The challenge of programming biology has been that we don't know biology." — Bart Ramsundar: Explaining why machine learning is helpful: it can learn from raw data where theory alone has been insufficient. "Anyone who can clone a repo can start really making a difference." — Bart Ramsundar: Describing how open source tools and accessible code are democratizing participation in biology. "You don't need a journal subscription to get Archive or to get the code, which is actually, that alone is kind of amazing." — Vijay Pandey: On the openness of modern ML research and how access has expanded beyond traditional academic gatekeeping.

Implications: Life sciences may become far more open, software-like, and participatory, but only if the field builds better representations, causality tools, and reproducible infrastructure. The winners will combine biological expertise with ML and embrace open collaboration.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast