The Cognitive Revolution
The Cognitive Revolution

Gemini Robotics – AI for the Physical World, with Keerthana Gopalakrishnan and Ted Xiao of Google DeepMind

In this engaging episode of the Cognitive Revolution, host Nathan Labenz welcomes guests Keerthana Gopalakrishnan and Ted Xiao to revisit significant advancements in robotics over the past year. Key themes discussed include the proliferation of new robotics companies, the emergence of humanoid robot

Featured Speakers

Nathan Labenz and Erik Torenberg Host

Topics Discussed

Episode Summary

Executive Summary: The episode argues that robotics is entering a GPT-3 to GPT-3.5-like phase: models are becoming broadly useful out of the box, embodied reasoning is improving, and fine-tuning with relatively little data can unlock strong task performance. Google DeepMind’s Gemini Robotics stack combines cloud-based high-level reasoning with fast on-device action control, but safety, data scaling, and embodiment constraints still limit deployment. The guests see humanoids as a major research frontier, while practical commercialization may arrive first in more controlled settings.

Main Topics: Robotics is entering a scaling era (Priority: 5/5): Kirthana Gopalakrishnan and Ted Zhao say the field has moved beyond simple tabletop demos into more realistic, commercially oriented embodied tasks. They frame the current moment as roughly between GPT-3 and GPT-3.5 for robotics, with increasing generality and out-of-the-box capability. Gemini Robotics architecture: embodied reasoning plus actions (Priority: 5/5): The release is split into Gemini Robotics ER, a cloud-run embodied reasoning model, and Gemini Robotics Actions, a distributed control model with high-frequency local actuation. The system is designed to merge high-level understanding with low-level motor control. Jagged capabilities and surprising dexterity (Priority: 4/5): The guests describe a frontier where robots can do impressive tasks like origami, Ziploc handling, tong use, and object manipulation, yet still fail on many tasks or require task-specific tuning. Dexterity has improved dramatically, but broad generalization remains uneven. Data scaling, imitation learning, and fine-tuning (Priority: 5/5): They emphasize that imitation learning works, including on more complex platforms like humanoids, and that even tens to hundreds of demonstrations can produce strong gains for narrow tasks. They also discuss the need for higher-quality, more diverse robot data and the promise of synthetic data. Safety and failure modes (Priority: 4/5): The conversation distinguishes between semantic safety (understanding harmful instructions) and operational safety (physical risk controls such as e-stops and force limits). Current systems are relatively stable and non-catastrophic in many cases, but safety is not yet at language-model maturity. Embodiment, humanoids, and deployment paths (Priority: 4/5): The guests debate whether humanoids or simpler platforms will reach society first. Humanoids are seen as a hard and inspiring research problem, while homes may be among the last environments to support safe, affordable deployment.

Key Arguments: Robotics is no longer in a 'tabletop demo' phase; the field has moved toward hard, real-world embodiment problems and commercialization. A useful metaphor for current robotics is GPT-3 to GPT-3.5, because models are becoming general and more usable out of the box without extensive task-specific fine-tuning. The key architectural insight is not rigid modularity versus end-to-end learning, but a hybrid system where cloud reasoning and local control cooperate through learned interfaces. Embodied reasoning improvements in the base model can lift downstream action performance because spatial understanding and physical commonsense are foundational to manipulation. The cloud model handles slower, higher-level reasoning, while the on-device model provides rapid, high-frequency motor control needed for dexterous physical interaction. Robotics failures are often stable and graceful rather than catastrophic, but safety must still be layered through semantic refusal, operational safeguards, and low-level control limits. Data quality and diversity matter as much as raw scale; generic repetitive robot tokens are not enough for AGI-level physical competence. Synthetic data from simulation and generative video is promising, but real-world data is still the gold standard for now. Humanoid robots are a major frontier, but their complexity may slow deployment; their main value today is as a research catalyst and testbed for physical AGI. Over time, more robotics capabilities will likely be upstreamed into general foundation models, reducing the need for specialized prompting and fine-tuning.

Data Points: Robotics maturity vs LLMs: Between GPT-3 and GPT-3.5 - Ted Zhao’s characterization of where robotics models currently stand technically. Control update rate: 250 milliseconds - Cloud-based embodied reasoning/planning updates in the Gemini Robotics stack. Low-level motor control rate: 50 cycles per second (50 Hz) - On-device action decoding for fast, reactive control. Fast adaptation examples: As few as 100 demonstrations - Some tasks can be improved substantially with very small numbers of examples. Fine-tuning range mentioned: 2,000 to 5,000 examples - Task-specific post-training used for stronger specialization in the paper. Safety benchmark performance: Around the 80% range - Referenced for common-sense harm avoidance / Asimov-style safety evaluation. Publicly available/publicly cited robot datasets: Tens of billions of tokens - Scale of current publicly available robotics datasets such as Open X-Embodiment. Long-term data ambition: 1 trillion tokens - A near-term scalable target discussed as desirable for robotics progress. Human-generated internet scale: Tens or hundreds of trillions of tokens - Used as a comparison for the scale robotics may eventually need to match.

Pivotal Quotes: "I think technically I would really put us somewhere between GPT-3 and 3.5." — Ted Zhao: Describing current robotics capability relative to the LLM timeline. "I don't think we have gotten the chat GPT yet." — Kirthana Gopalakrishnan: Explaining why robotics has not yet had a consumer-scale breakout moment. "I think if you think robotics is an AGI problem, then you would want to work with the best frontier model and add the action or the movement and physical reasoning as a capability on top." — Kirthana Gopalakrishnan: Arguing for foundation-model-centered robotics strategy.

Implications: Robotics is likely to get dramatically better soon, but deployment will lag capability because of safety, hardware, and data constraints. Expect more upstreamed capabilities, stronger base models, and early wins in controlled environments before homes and fully general humanoids.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution