The Cognitive Revolution
The Cognitive Revolution

Robotics Research Update, with Keerthana Gopalakrishnan and Ted Xiao of Google DeepMind

Google DeepMind researchers Keerthana Gopalakrishnan and Ted Xiao discuss their latest breakthroughs in AI robotics. Including models that enable robots to understand novel objects, learn from human demonstrations, and operate under ethical constraints. The conversation covers six groundbreaking pap

Featured Speakers

Nathan Labenz and Erik Torenberg HostTed Zhao GuestKirthana Gopalakrishnan Guest

Topics Discussed

Episode Summary

Executive Summary: Google DeepMind Robotics researchers Kirthana Gopalakrishnan and Ted Zhao describe a year of rapid progress toward general-purpose robots by combining internet-scale vision-language models, large multi-robot datasets, new prompting interfaces, and safety oversight. The conversation covers six projects that improve generalization across objects, embodiments, trajectories, feedback, and zero-shot control, while emphasizing that robust embodied AI still needs more data, better interfaces, and stronger safety/alignment methods.

Main Topics: RT2 and internet-scale multimodal transfer (Priority: 5/5): RT2 shows that co-training robot action data with web image-language data lets robots leverage broad world knowledge for object recognition and symbolic understanding, while preserving robot motion knowledge through mixed-batch training. Cross-embodiment generalization with RTX (Priority: 5/5): RTX aggregates data from many labs and robot morphologies to train one generalist policy across single-arm manipulators, demonstrating positive transfer and a new cross-robot benchmark for the field. Promptable robot learning with RT Trajectory (Priority: 4/5): RT Trajectory lets robots learn from line-drawing trajectory prompts and contact annotations, enabling finer control and faster adaptation from demonstrations than language-only instructions. Scalable oversight and robot constitution in AutoRT (Priority: 5/5): AutoRT explores how to run many robots in unseen environments with limited human supervision using LLM/VLM oversight plus a layered robot constitution of safety rules and task constraints. Teachability and faster learning from feedback (Priority: 4/5): Learning to Learn Faster focuses on iterative human feedback, nighttime distillation, and rollout-based planning to make robots learn from critique more efficiently and improve teachability. Zero-shot VLM guidance in Pivot (Priority: 3/5): Pivot shows that frozen vision-language models can choose robot motion directions from annotated images without special fine-tuning, suggesting VLMs already contain some embodied reasoning. State of robotics progress and open gaps (Priority: 5/5): The guests argue robotics is moving from a GPT-2 to GPT-3 era, but major gaps remain in data efficiency, dexterous manipulation, bimanual skills, action representations, and robust safety/alignment.

Key Arguments: Internet-scale foundation models materially raise robotics hit rates, turning many previously 20-30% reliable components into 60-70% reliable ones and accelerating the full research stack. Co-training with web data helps robots inherit broad concepts without catastrophic forgetting, but motion generalization still primarily comes from robot-specific demonstrations. Cross-embodiment training can outperform specialist models, implying robots are more alike conceptually than previously assumed even when morphologies differ greatly. New prompting interfaces matter: line drawings, contact annotations, and multimodal motion prompts can expose capabilities that language-only commands cannot. Scaling robot fleets requires scalable oversight; LLMs/VLMs can handle first-pass supervision and rule checks, but traditional robotics safety controls remain essential. Safety/alignment for robots is much harder than for text systems because failure can cause physical harm, so constitutions and hard engineering constraints are necessary. VLMs may already encode surprising physical priors and can sometimes guide robot action zero-shot, but learned fine-tuned policies still outperform them on reliability. General-purpose robotics is not solved yet; the next gains likely depend on better action interfaces, more dexterous representations, and more sample-efficient learning.

Data Points: RT1 robot demonstrations: 130,000 teleoperated expert demonstrations - In-house Google kitchen manipulation data used as the robot foundation for RT1 and referenced as the core robot dataset RT1 object diversity: ~17 objects - Limited office kitchen object set used in the RT1/RT2 data collection setup RTX robot embodiments: ~30 robot kinds / many labs - RTX combined datasets from many academic and industrial-style robot embodiments, starting from a restricted single-arm, single-camera setup AutoRT robot fleet: 50 unique robots and around 20 simultaneously - Scale of robots used to study autonomous data collection and oversight across unseen environments Reliability uplift from foundation models: 30% to 60-70% - Ted Zhao’s high-level claim about how internet-scale foundation models improve component-level reliability across robotics pipelines Generalization gap in RT2: Motion generalization remains limited - Robot-specific movement knowledge still comes mainly from robot data, not internet pretraining Human oversight ratio: About 5 robots per human - AutoRT description of humans maintaining line-of-sight control over multiple robots while LLMs/VLMs handled first-pass filtering Training cadence: Daytime interaction / nighttime distillation - Learning to Learn Faster uses a daily loop where user feedback is collected during deployment and distilled overnight Feedback optimization: Rollout search improves performance - Learning to Learn Faster used rollout-style planning over feedback trajectories to choose better learning paths Prompting iterations in Pivot: Multiple rounds of clustered annotations - Pivot repeatedly narrows motion choices using VLM-selected directions around annotated image points

Pivotal Quotes: "with the era of internet scale foundation models, things that used to work maybe 20-30% of the time are now working 60 to 70% of the time" — Ted Zhao: Explaining why robotics progress accelerated over the last year "people moved in the direction of thinking that all robotics, all robots are kind of similar. It's like it's only as different as English and Chinese or something" — Kirthana Gopalakrishnan: Describing the conceptual motivation behind cross-embodiment training in RTX "if no one's watching the robot, the rules could be watching the robot" — Kirthana Gopalakrishnan: Summarizing the rationale for AutoRT’s robot constitution and scalable oversight

Implications: Robotics is converging with foundation-model methods, but useful general-purpose robots will require better action interfaces, more dexterous data, and safety systems beyond prompting. Near-term gains may come from hybrid systems that combine large models with classical controls and human oversight.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution