Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

🔬 Training Transformers to solve 95% failure rate of Cancer Trials — Ron Alfa & Daniel Bear, Noetik

Today, we explain this piece of “clickbait” from our guest! TL;DR: 95% of cancer treatments fail to pass clinical trials, but it may be a matching problem — if we better understood what patients have which tumors which will respond to which treatments, success rates improve dramatically and millions

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The episode explores Noetic’s contrarian thesis that cancer drug failure is driven less by poor drug design than by poor patient selection. The founders describe building a human-tumor, multimodal data factory from scratch, training self-supervised foundation models on H&E, protein, spatial transcriptomics, and DNA, and using them to stratify patients, infer responders, guide trials, and support a $50M GSK licensing deal.

Main Topics: Noetic’s founding thesis: patient selection over drug design (Priority: 5/5): The company argues most oncology drugs fail because trials enroll the wrong patients, not because the molecules are inherently bad. Their goal is to model patient biology well enough to identify the right subpopulations for each therapy. Building a human-centered multimodal data pipeline (Priority: 5/5): Noetic generates its own data in-house from tumor sourcing through processing, with H&E, protein staining, spatial transcriptomics, and genotyping all captured from the same samples to create a rich training set. Why public datasets and cell lines are insufficient (Priority: 5/5): The speakers criticize immortalized cell lines and standard preclinical workflows as poor abstractions for patient biology, arguing that clinically useful models require intentional dataset design and human-derived samples. Virtual cell models as practical drug-development tools (Priority: 4/5): Rather than simulating every biochemical reaction in a cell, Noetic defines virtual cells as models that predict how patient biology changes under perturbation and help select targets, cohorts, and trial designs. Model validation, inference, and interpretability (Priority: 4/5): The models are used to recover known biology such as immune-hot vs immune-cold tumors, predict response clusters from H&E alone, and infer gene expression patterns that explain why some patients respond. Perturbation experiments and in silico humanized mouse models (Priority: 4/5): Noetic combines human-trained models with multiplexed CRISPR mouse tumor experiments to validate predictions and bridge human and mouse biology, including mapping mouse outputs into human gene space. Commercialization and industry adoption (Priority: 3/5): The team discusses a major GSK licensing deal as evidence that pharma increasingly wants broad model access across pipelines, not just one-off collaborations or molecule-level partnerships.

Key Arguments: Most oncology drug failures stem from incorrect patient selection rather than poor pharmacology or target choice. Clinically useful models must be trained on human patient data because cell lines and many animal models are too far from real patient biology. Rich multimodal datasets are necessary: H&E provides tissue context, protein stains provide cell-type context, and spatial transcriptomics provides molecular context. Self-supervised learning can uncover therapeutically relevant disease subtypes without relying on simplistic biomarkers like single mutations or single-gene signatures. H&E is a powerful inference substrate because it exists for nearly every clinical tumor sample and can be used retrospectively on old trials. Scaling data quality and diversity is essential; in biology, unlike language, the field is still far from having enough data to fully solve translational problems. Custom model architectures matter, especially when combining modalities and using long-context spatial information. Mouse perturbation systems remain useful if they are connected back to human biology through model-based translation and validation. Pharma companies want foundation models as reusable infrastructure across many programs, not just one collaboration. The practical goal of virtual-cell modeling is not perfect cellular simulation but better prediction of which patient should receive which drug.

Data Points: Cancer drug failure rate: 90%-95% - Used to motivate Noetic’s thesis that most drugs fail in the clinic. Approximate time to first trained model: ~18 months to 1.5 years - The team generated data for a long period before they could even train a model. Company timeline at the time of discussion: Year 4 - Ron says Noetic is in year four of the effort. Spatial transcriptomics assay runtime: ~2 weeks per run - A single machine run takes weeks, with two slides processed in a two-week run. Model training data scale: >100 million spatially resolved cells - Noetic says its paired multimodal dataset includes over a hundred million spatial transcriptomics cells. Comparative dataset scale: ~1.2 million images - ImageNet is cited as an analogy for a carefully curated, labeled dataset that enabled progress in vision. GSK licensing deal value: $50 million - The company announced a model licensing deal with GSK for OctoVC. Historical internal estimate for starting data generation budget: ~$10 million - Ron notes the early buildout may have been closer to $10M before later licensing value discussions. Immune checkpoint inhibitor response rate in lung cancer: 12% - Used as a validation example for whether the model can recover known responder subsets. Mouse system perturbation scale: ~100 knockouts per mouse/tumor panel - They describe multiplexed barcoded knockout experiments with roughly a hundred perturbations at once. Data reduction experiment impact: 10%-40% training data causes worse performance - They report performance drops and weaker cross-cancer generalization when training data is reduced.

Pivotal Quotes: "We all know the numbers that 90%, 95% of cancer drugs fail in the clinic. Why do they fail?" — Ron Alpa: Introduces the central thesis that patient selection is the core problem in oncology drug development. "There was no prior here that any of this would work. Like zero." — Ron Alpa: Describes the early stage of Noetic’s lab buildout and the uncertainty of the effort. "Our thesis is kind of that if you look at the data, a much richer kind of data... we're going to see that actually... what people thought was one subtype... is really three distinct subtypes of cancer." — Ron Alpa: Explains the belief that richer multimodal data reveals therapeutically relevant subtypes hidden by conventional pathology.

Implications: The episode suggests oncology AI will be won by teams that control high-quality human datasets, not just algorithms. If Noetic’s approach holds, H&E-first models could reshape trial design, diagnostics, and precision medicine across pharma.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast