Episode Summary
Executive Summary: Google DeepMind Robotics researchers Kirthana Gopalakrishnan and Ted Zhao describe a year of rapid progress toward general-purpose robots by combining internet-scale vision-language models, large multi-robot datasets, new prompting interfaces, and safety oversight. The conversation covers six projects that improve generalization across objects, embodiments, trajectories, feedback, and zero-shot control, while emphasizing that robust embodied AI still needs more data, better interfaces, and stronger safety/alignment methods.
Main Topics: RT2 and internet-scale multimodal transfer (Priority: 5/5): RT2 shows that co-training robot action data with web image-language data lets robots leverage broad world knowledge for object recognition and symbolic understanding, while preserving robot motion knowledge through mixed-batch training. Cross-embodiment generalization with RTX (Priority: 5/5): RTX aggregates data from many labs and robot morphologies to train one generalist policy across single-arm manipulators, demonstrating positive transfer and a new cross-robot benchmark for the field. Promptable robot learning with RT Trajectory (Priority: 4/5): RT Trajectory lets robots learn from line-drawing trajectory prompts and contact annotations, enabling finer control and faster adaptation from demonstrations than language-only instructions. Scalable oversight and robot constitution in AutoRT (Priority: 5/5): AutoRT explores how to run many robots in unseen environments with limited human supervision using LLM/VLM oversight plus a layered robot constitution of safety rules and task constraints. Teachability and faster learning from feedback (Priority: 4/5): Learning to Learn Faster focuses on iterative human feedback, nighttime distillation, and rollout-based planning to make robots learn from critique more efficiently and improve teachability. Zero-shot VLM guidance in Pivot (Priority: 3/5): Pivot shows that frozen vision-language models can choose robot motion directions from annotated images without special fine-tuning, suggesting VLMs already contain some embodied reasoning. State of robotics progress and open gaps (Priority: 5/5): The guests argue robotics is moving from a GPT-2 to GPT-3 era, but major gaps remain in data efficiency, dexterous manipulation, bimanual skills, action representations, and robust safety/alignment.
Key Arguments: Internet-scale foundation models materially raise robotics hit rates, turning many previously 20-30% reliable components into 60-70% reliable ones and accelerating the full research stack. Co-training with web data helps robots inherit broad concepts without catastrophic forgetting, but motion generalization still primarily comes from robot-specific demonstrations. Cross-embodiment training can outperform specialist models, implying robots are more alike conceptually than previously assumed even when morphologies differ greatly. New prompting interfaces matter: line drawings, contact annotations, and multimodal motion prompts can expose capabilities that language-only commands cannot. Scaling robot fleets requires scalable oversight; LLMs/VLMs can handle first-pass supervision and rule checks, but traditional robotics safety controls remain essential. Safety/alignment for robots is much harder than for text systems because failure can cause physical harm, so constitutions and hard engineering constraints are necessary. VLMs may already encode surprising physical priors and can sometimes guide robot action zero-shot, but learned fine-tuned policies still outperform them on reliability. General-purpose robotics is not solved yet; the next gains likely depend on better action interfaces, more dexterous representations, and more sample-efficient learning.
Data Points: RT1 robot demonstrations: 130,000 teleoperated expert demonstrations - In-house Google kitchen manipulation data used as the robot foundation for RT1 and referenced as the core robot dataset RT1 object diversity: ~17 objects - Limited office kitchen object set used in the RT1/RT2 data collection setup RTX robot embodiments: ~30 robot kinds / many labs - RTX combined datasets from many academic and industrial-style robot embodiments, starting from a restricted single-arm, single-camera setup AutoRT robot fleet: 50 unique robots and around 20 simultaneously - Scale of robots used to study autonomous data collection and oversight across unseen environments Reliability uplift from foundation models: 30% to 60-70% - Ted Zhao’s high-level claim about how internet-scale foundation models improve component-level reliability across robotics pipelines Generalization gap in RT2: Motion generalization remains limited - Robot-specific movement knowledge still comes mainly from robot data, not internet pretraining Human oversight ratio: About 5 robots per human - AutoRT description of humans maintaining line-of-sight control over multiple robots while LLMs/VLMs handled first-pass filtering Training cadence: Daytime interaction / nighttime distillation - Learning to Learn Faster uses a daily loop where user feedback is collected during deployment and distilled overnight Feedback optimization: Rollout search improves performance - Learning to Learn Faster used rollout-style planning over feedback trajectories to choose better learning paths Prompting iterations in Pivot: Multiple rounds of clustered annotations - Pivot repeatedly narrows motion choices using VLM-selected directions around annotated image points
Pivotal Quotes: "with the era of internet scale foundation models, things that used to work maybe 20-30% of the time are now working 60 to 70% of the time" — Ted Zhao: Explaining why robotics progress accelerated over the last year "people moved in the direction of thinking that all robotics, all robots are kind of similar. It's like it's only as different as English and Chinese or something" — Kirthana Gopalakrishnan: Describing the conceptual motivation behind cross-embodiment training in RTX "if no one's watching the robot, the rules could be watching the robot" — Kirthana Gopalakrishnan: Summarizing the rationale for AutoRT’s robot constitution and scalable oversight
Implications: Robotics is converging with foundation-model methods, but useful general-purpose robots will require better action interfaces, more dexterous data, and safety systems beyond prompting. Near-term gains may come from hybrid systems that combine large models with classical controls and human oversight.
From the Transcript
That happens is kind of interesting. I think my one-sentence explanation is that with the era of internet scale foundation models, things that used to work maybe 20-30% of the time are now working 60 to 70% of the time. And in robotics, right, as a very complicated, dynamic, engineered system with many pieces. In the past, if every small component of your entire system only worked 30% of the time, it would take many, many iterations to get a whole performance system working at scale. But now, when every single Part of the entire stack just works that much better from the research iteration process to the engineering scaling process to the data collection engines. I think you can really just see the pace increase when you just have many more successes and a much higher hit rate when you're going about and scaling up your research. So that's my one-sentence explanation, but definitely Kirtan, if you have any others. Yeah, like I said, I think last year we were, I said, what we lack is data. And so we have also tried to attack that problem from like different ends.
People used to think that all the robots are so different. All of their data is like so different. And people moved in the direction of thinking that all robots are kind of similar. It's only as different as English and Chinese or something. And the concepts are similar. It's just the manner of expression that's different. The data sets that resulted, even starting from those criteria, were so, so diverse, right? We have everything from like, you know, using like baby toys all the way to like industrial arms, you know, all the way to very dexterous cable routing robots. Even with this very limiting assumption, you still get so many different morphologies. If you train on VLMs and if you build on top of the knowledge of VLMs, then you can stitch a lot of concepts from the internet along with the emotions that you have in robotics data sets. You could, under the same initial conditions, just change the prompt a little bit, do some prompt engineering, and actually see qualitatively different behavior from the robot. Ultimately, if you want to do a lot of tasks and be
Robots of like these large fleets of robots go about in the real world, and we don't have as many people, we cannot scale the amount of people similarly, then we have to think about more creative ways of supervision and coordination and acting in the real world. And LLMs and VLMs are very smart already and they have understanding of zero-shot unseen scenarios already, and they can also reason about rules. So, if no one's watching the robot, the rules could be watching the robot. So, that's where. The loss came about. Like, how should a robot behave now that human attention is minimal? And what is also like the best, the most optimal way of distributing the limited amount of human supervision that we have? So, the way we used human supervision was like you had to speculate the supply of it and then distribute it for like teleoperation on tasks that the robot policies could not do or when they made mistakes.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co