Y Combinator Startup Podcast
Y Combinator Startup Podcast

Chelsea Finn: Building Robots That Can Do Anything

Chelsea Finn on June 17th, 2025 at AI Startup School in San Francisco.From MIT through her PhD at Berkeley, where she pioneered meta‑learning methods, and Google Brain, Chelsea Finn has built her career around teaching machines how to learn. Now an Assistant Professor at Stanford and co‑founder of P

Featured Speakers

Y Combinator Host

Topics Discussed

Episode Summary

Executive Summary: The talk argues that general-purpose robotics will outperform narrow, task-specific systems by combining large-scale real robot data, pretraining/post-training, and language-model-style architectures. The speaker details Physical Intelligence’s progress on laundry folding, home/mobile manipulation, and open-ended prompting, emphasizing that real data is essential, diversity improves generalization, and future gains will likely come from better post-training, RL, and infrastructure rather than only bigger models.

Main Topics: Why robotics needs foundation models (Priority: 5/5): The speaker contrasts today’s fragmented robotics industry—where each application requires a custom stack of hardware, software, and task logic—with the foundation-model approach of building one general-purpose robot model that can transfer across tasks and environments. Scale is necessary but not sufficient (Priority: 5/5): The talk explains that while scale matters for robotics as it does for language models, neither industrial automation data, YouTube videos, nor simulation alone solves the problem because each lacks either task diversity, embodiment alignment, or realism. Laundry folding as a proving ground for dexterity and long-horizon control (Priority: 5/5): A detailed case study shows how the team gradually trained robots to unfold, flatten, fold, and stack laundry, moving from simple single-shirt cases to crumpled multi-item baskets and improving through curated pretraining/post-training recipes. Pretraining + curated post-training as the key recipe (Priority: 5/5): The speaker argues that training on all data is not enough for complex tasks; instead, broad pretraining followed by fine-tuning on high-quality, consistent demonstrations unlocks more reliable performance and generalization. Generalization to new environments and robots (Priority: 4/5): The same model and recipe are extended to tidying, coffee making, and novel homes, showing that diverse pretraining data and better language-following can close much of the sim-to-real and environment gap. Open-ended language and hierarchical task decomposition (Priority: 4/5): The talk shows how synthetic prompts generated from robot videos can train a high-level planner to handle natural-language requests, corrections, and interjections, enabling the robot to follow more flexible user intent. Future infrastructure: RL, synthetic data, and deployment systems (Priority: 4/5): In Q&A, the speaker emphasizes the need for reinforcement learning, synthetic data for evaluation and online improvement, robust real-time inference infrastructure, and continued open-source and industry research.

Key Arguments: Robotics has been held back because solving one application often requires building an entire custom company and stack from scratch. A general-purpose robot model can be to physical tasks what foundation models are to language: reusable across tasks, environments, and embodiments. Scale alone does not solve robotics; data must also be diverse, realistic, and aligned with the target embodiment and task distribution. For complex dexterous tasks, the best results came from pretraining on broad robot data and post-training on curated, consistent demonstrations. Starting with simpler subproblems and progressively increasing difficulty was essential to making laundry folding work. Training on all available data without careful curation produced poor results on hard tasks; better quality and consistency improved performance substantially. A model pretrained across many tasks and robots can transfer to new tasks like cleaning, coffee making, and candle lighting without starting over. Diverse real-world home data materially improves performance in novel homes and reduces the generalization gap. Language-following is a major failure mode in robot policies and can be improved by preserving pretrained vision-language capabilities and tokenizing actions. Open-ended human commands can be handled by hierarchical policies trained with synthetic prompts relabeled from robot trajectories. Real robot data remains indispensable; simulation and synthetic data are useful supplements, especially for evaluation and online improvement, but not replacements. Reinforcement learning is likely to become a major part of post-training because it can increase success rate and efficiency through online data. Better infrastructure on both the robot side and the training side is a critical bottleneck for progress and deployment. There is strong industrial and investor interest because the technology is beginning to work in practice, not just in theory.

Data Points: Model size (early laundry policy): ~100 million parameters - Initial policy trained for folding a single-size, single-brand shirt using image-to-joint-position imitation learning. Control frequency: 50 Hz - Robot control loop for the early laundry folding policy. Company founding date: Mid-March 2024 - Physical Intelligence was founded shortly before the initial laundry experiments. Laundry task time (early reliable folding): ~20 minutes for 5 items - Early successful pretraining/post-training recipe could fold five clothing items but still slowly. Laundry task time after curation improvements: 12 minutes for 5 items - Improved curation strategy reduced time from about 20 minutes to 12 minutes for folding five items. Updated model size: 3 billion parameters - Open-source vision-language model (Polygemma) used as a pretrained backbone in later experiments. Action horizon: 50 actions / ~1 second - The diffusion head predicted a chunk of 50 future actions. Language-follow rate: 80% vs 20% - Tokenized-action / stopped-gradient recipe improved language following to 80% compared with 20% for the baseline approach. Novel home data scale: More than 100 unique rooms - Mobile manipulation data collected across homes, mock kitchens, and mock bedrooms for generalization experiments. Mobile manipulation share of pretraining mix: 2.4% - Tidying bedrooms and kitchens constituted a small fraction of the overall pretraining mixture. Generalization performance drop without static robot data: Below 60% - Excluding static robot data significantly reduced performance when evaluated in novel homes. Observed success rate in novel homes: Around 80% - Reported success rate for household generalization tasks, with several noted failure modes. Diversity effect: Performance rose to target-environment level - Increasing the number of represented homes improved performance to match training on the target environment.

Pivotal Quotes: "you essentially need to build an entire company around that application" — Speaker: Describing how today’s robotics stack is custom-built for each niche use case. "can enable any robot to do any task in any environment" — Speaker: Stating the goal of Physical Intelligence’s general-purpose foundation model. "scale is necessary ... but they're subordinate to actually solving the problem" — Speaker: Summarizing the view that data scale matters but is not sufficient on its own.

Implications: The field is shifting from bespoke robotics to reusable foundation models. Progress will depend on real data, diverse environments, better post-training, RL, and deployment infrastructure—opening the door to robots that are more adaptable, commercially viable, and user-directed.

🔓 Sign Up for Unlimited Episode Search

About Y Combinator Startup Podcast

We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.

View all episodes from Y Combinator Startup Podcast