Episode Summary
Executive Summary: Genesis Molecular AI argues that the hardest drug-discovery problems—especially protein-small molecule interaction and ADME optimization—are now becoming tractable through diffusion-based 3D structure prediction, synthetic-data generation, physics priors, and agentic workflows. Evan Feinberg and Sergei Udinov frame the company as a full-stack AI partner to pharma, focused on sub-angstrom accuracy, real program results, and faster design-make-test cycles that can produce better medicines.
Main Topics: Why diffusion became the right primitive for molecular AI (Priority: 5/5): The speakers contrast earlier GAN-based attempts with diffusion models, explaining that diffusion proved far more effective for 3D molecular and protein-ligand structure generation. They note that some of the most innovative diffusion work is now happening in 3D structure prediction. Protein-small molecule binding as the central bottleneck (Priority: 5/5): The conversation focuses on why predicting protein-ligand poses and binding is historically hard: vast search space, dynamic induced fit, and the need for atom-level precision. Genesis positions Pearl as a tool to make this problem computationally tractable and useful for downstream medicinal chemistry. Data, physics, and synthetic training sets (Priority: 5/5): Genesis says public structural data is limited, so it uses physics-based simulation and synthetic data to expand training sets. The model stack is designed to bake in physical priors and reduce the need for the model to relearn chemistry from scratch. One-angstrom resolution as the practical threshold (Priority: 5/5): A major thesis is that two-angstrom-level predictions are often insufficient for real drug design, because small errors can flip aromatic rings or ruin hydrogen-bond geometry. The company claims its focus on one-angstrom and sub-one-angstrom accuracy is what makes outputs useful for medicinal chemists and force-field methods. Beyond structure: potency, ADME, and multi-parameter optimization (Priority: 4/5): The guests stress that solving structure alone does not solve drug discovery. Genesis also targets potency, solubility, metabolic liabilities, toxicity, and other ADME properties, which are critical because many properties anti-correlate and require multi-objective optimization. Agentic drug discovery and lab-in-the-loop iteration (Priority: 4/5): The company is building an agentic platform (code-named Sapphire) in which LLM-like orchestration helps scientists use many tools and run design-make-test-analyze loops. The goal is not full automation of scientists, but much higher throughput and continuous learning from partner labs. Business model and pharma collaboration strategy (Priority: 4/5): Genesis positions itself as an AI company serving pharma and biotech rather than a traditional asset-centric biotech. The name change to Genesis Molecular AI reflects both that identity and its full-stack ambitions, with partnership examples including rapid progress toward development candidates and first-ever binders for hard targets.
Key Arguments: Diffusion was the correct generative primitive for molecular 3D problems after GANs failed to generalize to proteins and protein-ligand systems. Protein-small molecule interaction prediction is valuable only when accuracy reaches about one angstrom; coarser predictions can be misleading or unusable. Real drug discovery requires solving multiple correlated objectives simultaneously: potency, selectivity, solubility, PK, toxicity, and tissue distribution. Public structural datasets are too small to drive modern foundation models alone, so synthetic data and physics-based simulation are necessary. The company’s models are designed to be interoperable with medicinal chemistry and physics-based tools, not just generate abstract predictions. High-value drug discovery is not only first-in-class targets; improving existing clinical or approved molecules can produce large patient benefit. Agentic systems will augment rather than replace human scientists, making expert med chemists and CAD scientists more productive. Lab partnerships with fast feedback loops are essential because wet-lab results can be used directly for retraining and reinforcement learning. Benchmark-only progress is insufficient; models must be validated on real pharma programs and challenging out-of-distribution targets. The company believes the highest-leverage use of AI in healthcare is drug discovery and design, where success rates and value creation can be highest when biology is understood and molecules are well-optimized.
Data Points: Protein-coding genes: 20,000 - Evan cites the number of protein-coding genes to argue that many disease-causing targets remain to be addressed. Public crystal structures in PDB: ~200,000 - The team notes that the Protein Data Bank is relatively small compared with the needs of modern foundation-model training. Drug-like small molecules in chemical space: 10^60 - Used to emphasize the scale of search space for small-molecule discovery. Small-molecule share of FDA-approved drugs: 65% - Genesis says small molecules remain the largest approved modality and justify the company’s focus. Hydrogen-bond distance window: 2.7–3.3 Å - Illustrates why sub-angstrom accuracy matters in capturing meaningful molecular interactions. RMSD threshold discussed: 2 Å and below - The speakers argue that traditional RMSD benchmarks around 2 Å are insufficient for real drug-design use cases. ADME assays/properties referenced: 30+ - They describe a broad set of assays and properties needed for a molecule to be a viable drug. Genesis technical publication year: Last year / recent years - They reference a Pearl technical report, OpenBind results, and prior ADME/graph-ML publications as evidence of progress. Partner example: Insitx collaboration expansion - Used as a public example of progress from initial data to development-candidate work and first-ever binders.
Pivotal Quotes: "the right primitive to get created, and that turned out to be diffusion, which turned out to be a much more useful primitive for the space" — Evan Feinberg: Explaining why diffusion superseded GANs for molecular and protein-structure generation. "drug discovery really is a science of resolution" — Evan Feinberg: Arguing that atomic-level precision is necessary for practical medicinal chemistry and downstream physical methods. "we are definitely hiring" — Evan Feinberg: A direct call for engineers and scientists to join Genesis Molecular AI.
Implications: The field is moving from promising but fuzzy molecular AI to actionable, high-resolution systems that can materially change hit finding, optimization, and candidate design. For listeners, this suggests AI in drug discovery is entering a more practical, partnership-driven phase.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast