No Priors
No Priors

Virtual Cell Models, Tahoe-100 and Data for AI-in-Bio with Vevo Therapeutics and the Arc Institute

On this week’s episode of No Priors, Sarah Guo is joined by leading members of the teams at Vevo Therapeutics and the Arc Institute – Nima Alidoust, CEO/Co-Founder at Vevo Therapeutics; Johnny Yu, CSO/Co-Founder at Vevo Therapeutics; Patrick Hsu, CEO/Co-Founder at Arc Institute; Dave Burke, CTO at A

Topics Discussed

Episode Summary

Executive Summary: The episode centers on Tahoe 100, a 100M-cell perturbational single-cell RNA-seq dataset released by Vivo and ARC to accelerate “virtual cell” AI. The guests argue biology is now data-limited, not idea-limited: large, curated perturbation data can enable causal, generalizable models of cell state, improve target discovery, and eventually produce better drugs faster.

Main Topics: Tahoe 100 launch and dataset significance (Priority: 5/5): The guests explain Tahoe 100 as the world’s largest single-cell RNA-seq dataset and a landmark perturbational resource for training AI models in biology. Why virtual cell models are needed (Priority: 5/5): They argue protein language models capture structure and binding, but virtual cell models are needed to model context, dynamics, and cell-state responses to drugs or gene edits. Data quality, batch effects, and perturbational scale (Priority: 5/5): A major theme is that prior single-cell datasets were fragmented, observational, and batchy; Tahoe 100 and ARC’s atlas aim to provide cleaner, more diverse, causal data. Open sourcing and platform strategy (Priority: 4/5): Vivo and ARC explain why they are open sourcing the data: to move the field’s baseline upward, recruit the community, and operate as a lean platform company rather than a single-asset biotech. Scaling laws and AI progress in biology (Priority: 4/5): The speakers compare biology to LLM and image-model scaling, arguing that biology is at an early inflection point where more data plus better models should yield nonlinear gains. Drug discovery, clinical impact, and timelines (Priority: 4/5): They connect virtual cell models to target discovery, hit finding, and ultimately better therapeutics, while acknowledging clinical validation will take time. Industry competition and biotech business models (Priority: 3/5): The discussion closes on Chinese biotech competition, capital efficiency, CROs, and the need for smaller, faster, more innovative teams in US biotech.

Key Arguments: Biology is now data-limited at the cell-state level; single-cell perturbation datasets are the missing substrate for machine learning. Protein language models and structural biology are necessary but insufficient because they do not capture cell context, disease state, or response dynamics. Perturbational data shifts biology from correlation to causation by observing before/after states under gene or drug perturbations. Dataset quality matters as much as scale; prior public data were heavily biased toward healthy tissue and subject to batch and analytical effects. A large, curated, diverse atlas can reduce information redundancy and improve model generalization across tissues, diseases, and perturbations. Open sourcing the dataset is strategic: it speeds field-wide progress, invites external feedback, and lets a very small team leverage a much larger community. Virtual cell models could help identify the right target and the right chemical faster, reducing a major source of clinical failure. The field should be more hypothesis-free and ambitious now that sequencing, single-cell profiling, and compute are cheaper and more scalable. AI progress in biology should be judged by scaling laws and predictive performance, but clinical utility will lag model capability because wet-lab and trials are slow.

Data Points: Tahoe 100 dataset size: 100 million single-cell data points - Vivo/ARC’s perturbational single-cell RNA-seq release Public perturbational single-cell data before Tahoe: ~1–2 million single-cell data points - Estimate given for publicly available perturbational datasets worldwide ARC Virtual Cell Atlas observational data (SC Basecamp): ~230 million cells - Curated observational single-cell data assembled from public sources Combined atlas scale: ~330 million cells - 230M observational + 100M Tahoe 100 perturbational cells Pre-existing public human-cell collection: ~45–50 million cells - Estimate of human cells collated before SC Basecamp Drug treatments in Tahoe 100: 1,200 drug treatments - Depth of the dataset across perturbations Patient models in Tahoe 100: 50 different cancer models / patient-derived models - Dataset spans multiple patients and cancer contexts Drug-patient / drug-cell interactions: ~60,000 experiments - Operational scale described for Tahoe generation Team size to generate Tahoe 100: 4 people - Guests emphasized unusually high leverage and consistency Time to generate Tahoe 100: ~3 days - Described as the duration for the experimental execution once set up Current virtual cell predictive performance: ~10% prediction of DEGs - Best current models are described as weak at predicting differential gene expression Model token analog for Tahoe 100: ~200–300 billion tokens - Rough equivalence using 2,000–5,000 genes per cell Large language model scale reference: ~0.5–1 trillion tokens - Used as a rough benchmark for where scaling laws become apparent Evo 2 training data: 9.3 trillion nucleotides - Used to illustrate scaling and emergent biological capabilities Evo 2 BRCA1 variant performance: AUC ~0.94 - Example of zero-shot variant pathogenicity prediction Drug development failure rate: ~90% fail in clinical trials - Used to motivate better target selection and prediction

Pivotal Quotes: "The key is if you want a model that can learn about changes going on in the heart or in the brain or in the liver or in the bones, you need to be able to train across all of those different cell types." — Hani: Why broad, diverse perturbational data is needed for generalizable virtual cell models "The DNA lives in the ROM, the read-only memory... the RNA lives in the RAM, so it's like the working memory." — Patrick Su: Explaining the computer analogy for why transcriptomic context is central to modeling cell behavior "We think this is actually not only an additional data set for machine learning, we actually think it's the first data set that's going to enable machine learning in this space." — Johnny: Describing why Tahoe 100 may be a field-defining dataset

Implications: If the claims hold, biology shifts from small, hypothesis-heavy studies to large-scale, causal, model-driven discovery. That could improve target selection, cut trial-and-error, and eventually make drug discovery faster, cheaper, and more predictive.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors