The TWIML AI Podcast
The TWIML AI Podcast

π0: A Foundation Model for Robotics with Sergey Levine - #719

Today, we're joined by Sergey Levine, associate professor at UC Berkeley and co-founder of Physical Intelligence, to discuss π0 (pi-zero), a general-purpose robotic foundation model. We dig into the model architecture, which pairs a vision language model (VLM) with a diffusion-based action expe

Featured Speakers

Sergei Levine Guest

Topics Discussed

Episode Summary

Executive Summary: Sergei Levine explains Physical Intelligence’s goal of building general-purpose robotic foundation models, using Pi Zero and Pi Zero Fast as early steps. The conversation covers why robotics needs large, transferable models, how vision-language models plus diffusion-based action experts work, the importance of real-world data curation and post-training, and how better tokenization and open-sourcing could accelerate robot capability and adoption.

Main Topics: Mission: robotic foundation models for general-purpose robots (Priority: 5/5): Physical Intelligence aims to create a single adaptable model foundation for many robot tasks and embodiments, analogous to general-purpose language models like ChatGPT. Why robotics has lagged: data scarcity, generalization, robustness (Priority: 5/5): Levine argues robotics lacks internet-scale data, must handle physical edge cases, and needs more robust learning methods to succeed in the real world. Pi Zero architecture: VLM plus diffusion-based action expert (Priority: 5/5): Pi Zero combines a vision-language model with a smaller action expert that generates continuous trajectories, enabling dexterous control in real time. Recipe matters: pre-training and post-training data strategy (Priority: 5/5): The model’s success depends heavily on curated heterogeneous pre-training data and a smaller, high-quality post-training set, including mistakes and recoveries. FAST tokenization: better action compression for training speed (Priority: 4/5): FAST reframes actions in a frequency domain tokenization, improving information balance across tokens and dramatically improving training efficiency. Open-sourcing and ecosystem expansion (Priority: 4/5): The team released code, checkpoints, and demo models to encourage community experimentation and help refine future robotic foundation models. Future directions: complex instructions, reasoning, RL, and broader generalization (Priority: 4/5): Levine expects more progress on instruction following, task decomposition, test-time reasoning, and eventually reinforcement learning for robot models.

Key Arguments: Robotics needs foundation models because every new application currently requires enormous bespoke engineering and data collection. The biggest blockers in robotics are lack of large-scale data, poor generalization/common sense, and insufficient robustness/reliability. Vision-language models provide semantic priors, but dexterous control requires a separate mechanism for continuous spatial actions. Pi Zero is a starting point, not the final solution; it demonstrates the viability of the approach rather than full general-purpose robotics. Model quality depends as much on the training recipe—data mixture, curation, and post-training—as on architecture. Using only high-quality data is insufficient; robots also need mediocre and recovery data so they learn how to recover from mistakes. The pre-training/post-training split emerged organically once the team accumulated enough heterogeneous robot data and long training runs. FAST improves action representation by compressing trajectories in a frequency domain, making learning faster and more effective. Open-source release is meant to catalyze community experimentation, similar to what happened with language models. Future gains may come from better instruction following, more explicit reasoning, and reinforcement learning layered on top of the base foundation model.

Data Points: Pi Zero model size: 3.3 billion parameters - Referenced as the lightweight base model used for real-time robotic control. Pre-training data volume: ~10,000 hours - Large heterogeneous teleoperated dataset used to pre-train the model across many robots and tasks. Post-training data volume: 2 to 20 hours - Smaller, carefully curated high-quality dataset used to adapt the model to specific tasks. Action chunk rate: 50 time steps at 50 Hz - Pi Zero outputs roughly one second of future motion as a chunk, updated asynchronously. Inference cadence: about 3 times per second - Described as the approximate rate at which new action chunks override prior ones. FAST training speedup: about 4x faster - Tokenization in the FAST approach makes model training substantially more efficient than naive action tokenization. Droid dataset use: less than 10% of pre-training data - Used to increase embodiment diversity and help the model generalize across robot types. Robot embodiments in pre-training: 8 or 9 (pre-training set) vs about 20 (another dataset reference) - Used to illustrate that the foundation model spans more robot types than prior data collections. Aloha system cost: around $20K total - Cited as an example of a relatively affordable robot setup used in experiments. Per-arm cost: around $6K per arm - Approximate hardware cost for follower arms in the Aloha-style setup. Minimum practical setup estimate: around $15K - Estimate for a follower-arm-only setup to try the model. Action dimension limit: less than 32 dimensions - The model’s flexible action representation supports many embodiments as long as action dimensionality stays below this limit. Embodied Chain of Thought improvement: 50% better - Academic prior work from Levine’s lab reportedly improved performance with reasoning steps, albeit slowly.

Pivotal Quotes: "If we could have these general purpose models that can serve as the foundation for a huge range of applications, that would actually allow us to get robots to the next level." — Sergei Levine: He is motivating Physical Intelligence’s mission and the need for robotic foundation models. "The foundation model is not just about the model itself. It’s also about the recipe." — Sergei Levine: He emphasizes that data curation and training strategy are as important as architecture. "Using only high quality data. Only high quality... If you use only good data, it’s a bad idea." — Sergei Levine: He explains why recovery behaviors and error cases must be included in pre-training.

Implications: Robotics is moving from bespoke task systems toward reusable foundation models. If data curation, tokenization, and real-time control keep improving, general-purpose robots may become more practical, cheaper, and easier for others to adapt.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast