Episode Summary
Executive Summary: Brian He argues that biology’s biggest bottleneck is causal understanding, not binder design, and that AI can help by learning from both observational and interventional data. He discusses three projects: architecture search for better hybrid models, Evo’s surprising biological concept learning from DNA alone, and structure-informed antibody design that improved binding up to 25x.
Main Topics: Grand challenges in AI for biology (Priority: 5/5): The conversation frames biology as two core problems: identifying causal mechanisms in disease and designing precise interventions. Brian emphasizes that most drug failures stem from targeting the wrong biological mechanism, not from weak binders. Mechanistic Design and Scaling of Hybrid Architectures (MAD) (Priority: 4/5): A semi-automated framework for composing hybrid ML architectures from primitives like attention and state-space layers. The paper shows micro-skill performance predicts larger-scale performance and hybrid models often outperform pure architectures. Evo and DNA-only sequence modeling (Priority: 5/5): Evo is a long-context hybrid model trained only on DNA that nonetheless learns higher-level biological regularities. It predicts gene essentiality, suggests CRISPR variants, and indicates that raw evolutionary sequence data contains rich latent biological concepts. Biosecurity and responsible release (Priority: 4/5): The discussion covers exclusions made in model training, the need for evals on harmful sequences, and the view that generative biology tools are more useful for defense and medicine than for offense. Structure-informed protein and antibody engineering (Priority: 5/5): Using a structure-conditioned model trained on single proteins, Brian’s team generalizes to multi-chain complexes and proposes mutations that can improve antibody-target binding by large margins. Interpretability and future biology models (Priority: 3/5): Brian expects sparse feature or concept-level interpretability to help identify biological concepts inside models, enabling low-shot discovery, classification, and design in underexplored systems. ARC Institute and scientific execution (Priority: 2/5): Brian describes ARC as a highly collaborative, talent-dense environment enabling rapid progress on frontier biology and AI research.
Key Arguments: Most drug failures occur because the target is wrong, not because the binder is poor; causal mechanism discovery is the central bottleneck. Observational biology data is abundant but insufficient for causality because spurious correlations dominate; interventional data is needed. Hybrid architectures can outperform pure attention or pure state-space models, and small-scale micro-skill benchmarks can predict larger-scale results. Smaller, more data-trained models may be especially valuable in biology because inference has to be usable by wet-lab researchers without giant compute resources. DNA sequence alone can encode enough evolutionary signal for models to infer higher-order concepts like gene essentiality and conserved biological function. Evo’s gene essentiality results suggest the model is doing more than memorizing conservation; it captures context-dependent biological importance. Generative models may be able to produce biologically meaningful variants beyond naturally observed sequences, including novel CRISPR-like systems. Safety should focus on exclusion, evaluation, and responsible standards; the bigger societal value is on defense, vaccines, and therapeutics rather than offense. Structure-informed models trained on single proteins can generalize to complexes, showing strong out-of-distribution capability in protein design. Biology often needs low-throughput, high-confidence design; AI is most valuable when it raises the hit rate of a handful of expensive wet-lab tests.
Data Points: Evo context length: 131 kilobases - Brian says Evo can fit genomes up to about 131,000 nucleotides in context. Bacteriophage genome size example: ~40 kilobases - A bacteriophage genome used in the gene-essentiality experiment fit within Evo’s context window. Insertion size for knockout test: 3 letters - They inserted a premature stop mutation as a small three-letter change to test gene essentiality. Model layers using attention in Evo: 3 out of 32 layers - Brian says Evo uses only a small fraction of attention layers, roughly one in ten. Binding improvement: up to 25x - The antibody design work found variants that bind their targets up to 25 times better than natural counterparts. Experimental rounds in antibody evolution: 2 rounds - They used a first round of model-ranked mutations and a second round combining beneficial mutations. Single-mutation screening scale: sequence length times number of amino acids - Brian describes scoring all possible single amino-acid changes across the antibody sequence. Typical lab throughput for CRISPR variants: 10 to 15 - Brian notes biological assays may only practically test around 10–15 candidates in a reasonable time and budget. Training data exclusion: eukaryotic viruses excluded - For safety, sequences from viruses that infect eukaryotes were left out of the Evo training set. Professor tenure at ARC: 8 months - Brian says he has been a professor for only about eight months at the time of recording.
Pivotal Quotes: "The reason why most drugs fail out in the clinic is because you just went after the wrong target to begin with." — Brian He: Explaining why causal mechanism discovery matters more than binder optimization in drug development. "The model has learned something. How do we now start to interpret this model and connect it to actual biological discoveries?" — Brian He: On the central challenge of translating model representations into usable biological insight. "It’s much harder to create or to ameliorate bad biology than it is to destroy." — Brian He: His core argument for why AI-enabled biology is more valuable for defense and medicine than for offense.
Implications: AI for biology is shifting from pattern recognition to causal discovery and design. The near-term winners will likely combine better models, better benchmarks, safer release practices, and tightly coupled wet-lab validation.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co