Y Combinator Startup Podcast
Y Combinator Startup Podcast

Robot-Use Agents: Why General-Purpose Models May Win in Robotics

One of the biggest surprises in AI over the last few years has been how well coding agents generalize beyond software. In a recent essay, MIT professor Philip Isola argued that we may be entering the era of robot-use agents: general-purpose models that can control different robots, write policies, a

Featured Speakers

Y Combinator Host

Topics Discussed

Episode Summary

Executive Summary: The discussion argues that recent advances in large language models are making robot control far more capable, especially when models are paired with harnesses, tools, and code-based policies. The guests emphasize transfer from web/code/computer-use data into robotics, the importance of in-context learning and skill distillation, and why frontier models may soon enable general-purpose robot agents.

Main Topics: LLMs as robot controllers (Priority: 5/5): The conversation centers on how frontier language models can directly control robots, either through action outputs, tool calls, or code generation, and why this is suddenly working better than earlier approaches. RT-2 and early vision-language-action models (Priority: 5/5): The guests review early robotics foundation work, especially RT-2, as a milestone showing that pretrained language models can improve robotic control when fine-tuned to output actions. Code as policies and program induction (Priority: 5/5): They connect coding agents, Voyager, and code-as-policies work to robotics, arguing that code is a powerful way to compress thinking into reusable skills and create new tools on the fly. Bitter lesson and data transfer across domains (Priority: 5/5): A major theme is that general-purpose models benefit from scaling and diverse pretraining data more than specialized robotic architectures, with language/computer-use data helping robotics. In-context learning vs weight updates (Priority: 4/5): The panel compares learning in context, tool use, LoRA, and full fine-tuning/RL, arguing that context is cheap and flexible but saturates, so compression into skills and weights is needed. Harnesses, skills, and latency reduction (Priority: 4/5): The startups frame their systems as harnesses that wrap models, store skills, and consolidate repeated behavior so robots can act faster, more reliably, and with lower latency. Platonic representations and multimodal convergence (Priority: 4/5): They discuss the idea that strong models across modalities may share similar world representations, implying that improving language models may also improve robot intelligence.

Key Arguments: Robotics is beginning to benefit from the same scaling dynamics that made coding agents work, because stronger general-purpose models can transfer across domains. RT-2 was important because it used web text/image pretraining and mapped model outputs into robot actions, proving pretraining helps robotics. The main bottleneck is increasingly data and harness design, not just architecture; diverse modalities like coding, computer-use, egocentric video, and CAD may unlock better robot control. Code is a compact, reusable form of reasoning that can generate tools and policies, making it a strong interface for robot action. In-context learning can rapidly improve performance on held-out tasks without SGD, but it saturates quickly and cannot replace weight updates indefinitely. A practical robotics system should combine fast in-context adaptation with slower consolidation into skills, policies, or weights during a sleep/reflection phase. Computer-use data is valuable for robotics because it teaches spatial concepts like left/right, top/down, and manipulation of 2D/3D environments, which transfer to physical control. The guests believe frontier models may enable general-purpose robots within about two years, though latency and skill management remain major challenges.

Data Points: Latency improvement for Fable-class LLMs: around 2x per month - Mentioned as a trend that could enable real-time robot control by the end of the year. Context saturation for in-context learning: about 20–40 examples - A cited experiment showed performance improvements saturating quickly after a small number of examples. Training-context limit example: 100,000 training context / 50,000 effective context - Used to illustrate that performance degrades once context length exceeds what the model can attend over reliably. General-purpose robot timeline: within the next 2 years or even earlier - Consensus claim from frontier labs and robotics foundation model companies.

Pivotal Quotes: "If you give the agents or if you give the AI model more autonomy and you unshackle it a bit more and give it more resources, it can actually do a lot of the things that we fine-tune it to do." — Jay: On the bitter lesson and why general-purpose models may surpass specialized robotics models. "We want to benefit from all kinds of data. We want to pour in computer-use data into our robot models. You want to pour encoding data into our robot models." — Francois: On transferring diverse modalities into robotics through a general-purpose model. "There's some consensus within the Frontier Labs and also in the Robotics Foundation models companies that we will have general purpose robots within the next two years or even earlier." — Jay: On the expected near-term arrival of broadly capable robots.

Implications: Robotics is shifting toward general-purpose, multimodal agents built on frontier LLMs plus harnesses and skill distillation. Expect faster progress, more reusable robot behaviors, and a near-term push to solve latency, memory, and deployment at scale.

🔓 Sign Up for Unlimited Episode Search

About Y Combinator Startup Podcast

We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.

View all episodes from Y Combinator Startup Podcast