Episode Summary
Executive Summary: The conversation centers on BioHub’s ESMC protein language model and its new “world model” approach to programmable biology. Alex Reeves argues that scaling sequence data, especially metagenomics, unlocks emergent biological representations and better protein design without heavy priors. The discussion expands to interpretability, antibody design, interactome mapping, and BioHub’s broader vision of building open, feedback-driven AI and experimental infrastructure for cellular and disease biology.
Main Topics: ESMC as a world model for protein biology (Priority: 5/5): ESMC is presented as a predictive model that searches latent biological space to design proteins, binders, antibodies, and SCFVs, rather than relying on hand-built priors. Scaling laws, bitter lesson, and data-driven biology (Priority: 5/5): Reeves emphasizes that protein biology follows scaling behavior: larger models and larger, more diverse datasets produce emergent capabilities, validating a bitter-lesson-style approach. Metagenomics and the jump from ESM2 to ESMC (Priority: 5/5): The key breakthrough behind ESMC was adding massive metagenomic sequence diversity, which removed diminishing returns seen in ESM2 and improved representational fidelity. Mechanistic interpretability and sparse autoencoders (Priority: 4/5): Sparse autoencoders reveal a hierarchical feature space inside ESMC that mirrors known biological concepts, from biochemical motifs to higher-level functional themes. Programmable biology and therapeutic design (Priority: 4/5): The model is used to design protein binders and antibodies, with particular excitement around SCFVs and the possibility of reformatted antibody therapeutics. Virtual biology, cellular modeling, and feedback loops (Priority: 5/5): BioHub’s broader mission is to build models of cells and physiology, paired with experimental platforms that generate feedback and enable iterative scientific discovery. Open science infrastructure and BioHub’s mission (Priority: 4/5): The episode frames BioHub as an open, philanthropic scientific institution building measurement technologies, atlases, and AI tools to accelerate disease understanding and prevention.
Key Arguments: Protein biology can be learned through scaling over huge sequence corpora, similar to language modeling, because sequence patterns encode structural and functional constraints from evolution. Metagenomic data was the decisive improvement over earlier datasets like UniRef, because it adds massive diversity from many ecological niches and reduces the data limitation of earlier models. Evolving proteins are not independent sequences; amino-acid choices are constrained by structure and function, so language-model training can recover latent biological variables. Sparse autoencoders show that the model organizes biology into a hierarchy resembling the reductive framework biologists developed experimentally, but learned without explicit supervision. ESMC supports programmable biology through world-model search: instead of imposing structure priors, it can search latent space for molecules that satisfy design criteria. The model can design therapeutically relevant antibodies and SCFVs with promising affinity, suggesting utility beyond simpler mini-binders. A future “virtual biology” stack requires three things: scale in data generation, predictive digital representations, and feedback from experiments to update the models. Cell-level modeling is the next major bottleneck, and it will require new data modalities, spatial biology, perturbation biology, and advanced experimental technologies. BioHub’s role is to build the infrastructure for this new paradigm, not to become a drug company, but to create open tools that accelerate science broadly.
Data Points: Protein sequences in BioHub atlas: 6.8 billion - Non-redundant proteins assembled across the world’s largest protein sequence databases Predicted structures in atlas: 1.1 billion - Structures resolved/predicted for clustered representative proteins ESMC model sizes: 300M / 600M / 6B parameters - Protein language model family used for representation analysis and downstream tasks ESM2 scale comparison: ~1B to 10B parameters - Previous generation showed improvements but diminishing returns on earlier datasets BioHub Virtual Biology Initiative internal funding: $400 million - Announced commitment for data creation and technology development BioHub Virtual Biology Initiative external catalysis funding: $100 million - Funding to catalyze outside efforts to generate biological data Estimated current cell atlas scale: ~1 billion cells - Scale of largest current cell biology efforts referenced in the discussion Sequence clustering threshold: 70% sequence identity - Used to cluster the 6.8B protein database and choose structural representatives Antibody share of therapeutics: ~25% of new drugs - Used to emphasize the importance of antibodies as a modality Training-data scale change: Order-of-magnitude more sequences - ESMC benefits from substantially more data than ESM2 via metagenomics Prior model data source: UniRef - Described as the gold-standard curated protein sequence dataset used for ESM2 Largest scale referenced for future data: ~100 billion sequences - Rough estimate of the broader protein sequence universe discussed as still untapped
Pivotal Quotes: "I believe in scaling laws." — Alex Reeves: Explaining the philosophical basis for ESM’s approach to protein modeling "The idea is basically you have a predictive model and you're going to search the world model to find protein molecules that satisfy kind of whatever design criteria that you have." — Alex Reeves: Describing ESMC’s programmable biology and protein design workflow "We want models that can serve as oracles for the biology. They can predict an experiment that you haven't done." — Alex Reeves: Defining the standard BioHub wants for virtual biology and cell modeling
Implications: The episode suggests protein design is entering a data-scaling era where open foundation models can accelerate discovery, therapeutic engineering, and eventually cell-level biology. It also implies the next breakthroughs will depend less on priors and more on massive, diverse datasets plus experimental feedback loops.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast