Episode Summary
Executive Summary: Zera Therapeutics is building an AI-native drug discovery platform centered on “virtual cell” models. The discussion explains why causal, genome-wide perturbation data—not just observational single-cell data—are needed to predict cellular responses, how Zera’s high-throughput Perturb-seq experiments generate that data, and why their diffusion-based Excel model outperforms linear baselines and generalizes across unseen cell types and primary cells.
Main Topics: Zera’s end-to-end AI drug discovery thesis (Priority: 5/5): The founders describe Zera as an AI-enabled therapeutics company spanning target ID, protein design, cellular modeling, and patient representation, aiming to accelerate discovery and improve clinical success rates. Virtual cell as causal biology prediction (Priority: 5/5): They distinguish virtual cells from older mechanistic and foundation-model approaches, arguing the goal is to predict perturbation outcomes, dynamics, and context-specific cell behavior rather than only static representations. Why causal data beats observational data (Priority: 5/5): The guests argue that descriptive datasets like large single-cell atlases are excellent for representation learning but insufficient for causal prediction; genome-wide perturbation datasets are required to learn intervention effects. High-throughput Perturb-seq data generation (Priority: 5/5): Zera’s lab workflow uses pooled CRISPR perturbations plus single-cell RNA-seq to systematically knock down genes at scale, creating large causal datasets across many cell types and contexts. Excel model architecture and priors (Priority: 4/5): Excel shifts from autoregressive gene-token modeling to diffusion-based generation and incorporates multiple biological priors (literature, PPI, cancer essentiality, morphology, cell-type embeddings) to improve generalization and interpretability. Generalization across contexts (Priority: 5/5): The paper’s key claim is that Excel predicts perturbation responses better than linear baselines and can transfer across activated vs resting T cells, held-out cell types, and cell lines to primary cells. Open science, industry scale, and the future of biology AI (Priority: 4/5): The speakers emphasize open-source models/data, argue that academia should focus on high-taste innovation and novel biology, and identify protein measurement and longitudinal single-cell measurement as the next major bottlenecks.
Key Arguments: Observational single-cell data are underpowered for causality; many causal graphs can fit the same descriptive data, so perturbation datasets are necessary to learn true intervention effects. High-throughput Perturb-seq creates causal, genome-wide training data by combining pooled CRISPR perturbations with single-cell RNA-seq, enabling scalable intervention-response mapping. Excel’s diffusion language model better matches the unordered structure of expression matrices than autoregressive tokenization and improves perturbation prediction on harder generalization tasks. Adding diverse biological priors helps performance and interpretability, but data quality and scale contribute most to the model’s success. Generalization across unseen contexts—cell types, activation states, and primary cells—is the real test for virtual cell models, not just low MAE on small benchmarks. AI models should not replace wet-lab work; they should generate better hypotheses for complex systems where exhaustive experimentation is impossible. Academia still matters because it drives foundational innovation, teaches scientific taste, and has access to unique clinical/healthcare datasets, while industry is better positioned to industrialize and scale data generation.
Data Points: Human genes per cell: ~20,000 - Used to explain the scope of perturbation and transcriptomic readout in a cell Genes expressed in a typical human cell: ~4,000–5,000 - Described as the active subset in a given cell state Protein structure curation history: ~70 years - Cited as the long-term accumulation that enabled protein-model breakthroughs Single-cell atlas size: 33+ million cells - Referenced for CellxGene-like observational datasets Genome-wide perturbation campaign count: 7 screens - The dataset integrated multiple perturbation campaigns across contexts Biological contexts in latest dataset: 16 cell types - Used to emphasize diversity and richness of training data Filtered high-quality cells: 25 million cells - Number retained after stringent QC from much larger starting material Model size: 4.5 billion parameters - Excel model scale mentioned during discussion Clinical trial success rate: 5–10% - Phase 3 success rate cited as a major motivation for better prediction models Disease burden: 90% of disease has no cure - Used to motivate the need for better therapeutic discovery
Pivotal Quotes: "This is the first time that someone can put together not just one perturbed seq, but seven genome-wide perturbation campaigns together." — Speaker on the podcast: Explaining why the new dataset felt like a breakthrough to biologists "If we cannot describe it, let’s learn it." — Si Chu: Summarizing the virtual cell 2.0 philosophy of data-driven biological modeling "What really blew my mind away is when I saw the model make prediction... it is visually very clear to see that Excel prediction is much more similar to ground truth than the linear baseline." — Podcast guest/commentator: Describing the visual comparison between baseline, ground truth, and Excel predictions
Implications: The takeaway is that AI-for-biology progress will depend less on bigger generic models and more on high-quality causal data, cross-context validation, and open collaboration. If this works, it could compress drug discovery cycles and improve patient matching.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast