Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

🔬 The Lab of the Future Should Feel Like a Data Center — Andy Beam & Rafa Gómez-Bombarelli, Lila Sciences

Imagine a dark warehouse. Racks and racks of devices with wires, tubes, and electronics sticking out. The next AI data center? No. This is Lila Sciences‘ dream for the future of science. A dark warehouse full of AI-guided robotics and lab equipment, cranking out new experiments 24/7, building toward

Featured Speakers

Latent.Space HostAndy Beam Guest

Episode Summary

Executive Summary: Lila Science argues that science can be scaled like AI by turning experiments into verifiable data-generation loops. The company is building an AI-driven “science factory” that combines frontier models, lab automation, and human-in-the-loop verification to accelerate discovery across biology, chemistry, and materials, with an emphasis on generality, iteration speed, and real-world validation rather than narrow automation or a biotech-only strategy.

Main Topics: Lila’s core thesis: science as a scaling problem (Priority: 5/5): The guests frame Lila as applying the bitter lesson to science: general, scalable methods beat narrow ones. They argue the next frontier is replacing internet text with experimentally generated scientific data and using science itself as a verifier for model post-training. AI science factories and the lab as a data center (Priority: 5/5): They describe Lila’s platform as an AI science factory: a lab designed for continuous data generation, where instruments, humans, and software form an API-like system that feeds reasoning models through verified experimental feedback. Automation, flexibility, and hybrid human-robot workflows (Priority: 4/5): Rather than pursuing full automation everywhere, Lila prioritizes flexibility and generalizability. Some steps are robotic, some remain human-performed, and the key is that every action is exposed as a controllable, legible tool call to the model. Scientific RL, verification, and safety (Priority: 4/5): The discussion emphasizes reinforcement learning with experimental feedback, the risks of reward hacking/pathologies, and the need for strong lab safety, data security, and scientific rigor because AI-driven science can produce genuinely dangerous or misleading outcomes. Cross-domain generalization across bio, chemistry, and materials (Priority: 5/5): The team says a broad reasoning model trained on experimentally verified traces across many scientific domains outperforms domain-specific models, with transfer observed between areas like drug discovery, electrocatalysis, sorption, and materials formulation. Commercial model: virtual startups and platform partnerships (Priority: 4/5): Lila is positioned not as a biotech asset company but as a platform that can run many scientific programs for partners, effectively serving as a ‘Claude Code for science’ or a virtual startup engine that accelerates R&D for customers. Materials and bio as difficult but strategic frontiers (Priority: 4/5): The speakers compare biology and materials as hard, long-horizon domains with different bottlenecks. Biology has stronger commercialization pathways; materials has harder simulation and scale-up problems but huge implications for energy, sustainability, and industrial applications.

Key Arguments: General, scalable methods outperform narrow ones, especially when paired with large amounts of high-quality feedback. Internet text is a finite resource; science can become the next large-scale data source through experiments and verifiers. Experiments should be treated as data-generation and verification loops, not just one-off lab work. Lila is not trying to maximize raw automation; it is maximizing useful scientific tokens and experimental flexibility. Human technicians remain part of the workflow when that is faster or safer than automating every step. Scientifically verified reasoning traces are extremely valuable training data and likely rare in pretraining corpora. A broad scientific model can transfer across domains better than siloed models because human scientists already reason across domains using shared language and tools. The company believes iterative rounds of experimentation matter more than broad but noisy multiplexing in many settings. Safety and scientific rigor must be preserved because capability jumps can be sudden and unexpected. Commercial value comes from platform leverage: one system can support many programs, customers, and eventual product pathways.

Data Points: Scientific reasoning corpus size: 10 trillion tokens - Lila says it has assembled a reasoning dataset of experimentally verified scientific tokens across life sciences, chemistry, and materials. Zero-shot model performance on gene-editing expression protocols: ~80% - Andy said the model achieved about 80% zero-shot performance on certain expression-protocol tasks, versus humans at 0% zero-shot. Human zero-shot performance on same task: 0% - Used as a contrast point for model capability on a specific protocol-design benchmark. In vivo CAR-T project turnaround: ~6 months - A small team progressed from idea to non-human primate data in about six months. Expression improvement over references: ~10x - They said their untranslated-region designs achieved about 10x expression compared with Moderna and Pfizer references. B-cell depletion outcome: Better than Capstan data - They claimed their non-human primate in vivo CAR-T results outperformed recently public Capstan benchmarks on several characteristics. Capstan acquisition value: $2.1 billion - Used as a reference point for the market value of in vivo CAR-T progress. Traditional CAR-T infusion cost: ~$400,000 per infusion - Mentioned while explaining why in vivo CAR-T delivery could be transformative. Clinical trial approval rate: ~5-8% - Used to emphasize how hard downstream translation remains even after discovery. BET sorption assay speed: ~1 day per sample - Referenced as a slow conventional measurement that Lila is trying to replace with faster proxy readouts. Sorption proxy assay speedup: ~96 samples in ~1 hour - They said their 96-well proxy approach is about 2,500x faster than the traditional method. Training efficiency on GPUs: ~5-6% MFU - Andy cited mean flop utilization for RL training as only a small fraction of peak GPU throughput. Office/facility size: 100,000 sq ft - They mentioned an upcoming Cambridge facility of this size. Current San Francisco office headcount: ~20-30 people - Closing remarks about the 181 Fremont office and hiring plans.

Pivotal Quotes: "We are all in on the bitter lesson and scale." — Andy Beam: Summarizing Lila’s philosophy that general, scalable methods beat narrow ones in scientific AI. "We think that the lab of the future should feel like a data center." — Andy Beam: Describing the desired end-state for automated, high-throughput scientific infrastructure. "Science is as an infinite token generator to train models at scale." — Andy Beam: Explaining why science can supply the next large-scale training signal after the internet.

Implications: If Lila’s approach works, scientific discovery could shift from artisanal lab work to scalable, verifiable model training loops. That could compress R&D timelines, improve translation odds, and create a new platform category spanning biotech, chemistry, and materials.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast