The TWIML AI Podcast
The TWIML AI Podcast

Waymo's Foundation Model for Autonomous Driving with Drago Anguelov - #725

Today, we're joined by Drago Anguelov, head of AI foundations at Waymo, for a deep dive into the role of foundation models in autonomous driving. Drago shares how Waymo is leveraging large-scale machine learning, including vision-language models and generative AI techniques to improve perceptio

Featured Speakers

Dragomir Anguilov Guest

Topics Discussed

Episode Summary

Executive Summary: Waymo’s research head Dragomir Anguelov described how foundation models, multimodal learning, and generative AI are being adapted to autonomous driving. He emphasized Waymo’s scale and safety progress, but argued that driving still requires specialized handling of 3D spatial reasoning, long temporal memory, sensor fusion, simulator realism, and hallucination control.

Main Topics: Waymo’s scale and operational expansion (Priority: 5/5): Waymo has moved from early Phoenix testing to a multi-city, fully autonomous consumer service across San Francisco, Phoenix, Los Angeles, and Austin, serving customers at meaningful scale. Safety performance and public validation (Priority: 5/5): Drago highlighted Waymo’s safety reporting, third-party insurance validation, and the company’s mission to be the world’s most trusted driver, using crash and injury statistics as evidence. Foundation models and VLMs in autonomy (Priority: 5/5): Waymo is adapting multimodal foundation models for driving, leveraging world knowledge from internet-scale pretraining while fine-tuning for perception, planning, and scene understanding tasks. Why autonomous driving is different from generic VLMs (Priority: 5/5): Driving requires strong 3D spatial awareness, long memory over scenes, multi-sensor input (LiDAR/radar/cameras), and robust hallucination mitigation—capabilities not native to standard VLMs. World models, simulation, and validation (Priority: 4/5): The team sees predictive world models as useful both for training and for building better simulators, especially because validation at AV scale depends on realistic, generative simulation and rare-case coverage. End-to-end vs modular autonomy stacks (Priority: 4/5): Waymo favors practical system design: fewer, larger components and selective end-to-end training, but not at the expense of testability, controllability, and fast debugging. Community challenges and research collaboration (Priority: 3/5): Waymo continues releasing datasets and hosting academic challenges focused on camera-only end-to-end driving, simulated agents, traffic generation, and interaction modeling.

Key Arguments: Foundation models are valuable because they bring broad world knowledge that can reduce manual labeling and improve generalization to new concepts and scenarios. Autonomous driving is not a vanilla VLM problem: it needs 3D understanding, temporal context, multi-sensor fusion, and protection against hallucinations. Scaling models in the data center can make them teachers for onboard systems and accelerate adaptation to new cities, countries, and sensor configurations. The best way to scale validation is to improve simulation with machine learning, because real-world testing alone cannot cover the long tail of rare and unsafe scenarios. End-to-end learning is attractive, but full sensors-to-controls systems are hard to test, hard to explain, and hard to patch safely at deployment scale. Waymo combines imitation learning, rule-based constraints, expert judgment, and testing harnesses because no single method is sufficient for safe AV release.

Data Points: Autonomous trips per week: 200,000+ - Waymo says it is serving paying customers at scale across four markets. Markets served: 4 - San Francisco, Phoenix, Los Angeles, and Austin. Autonomous miles per week: More than 1 million - Waymo’s fleet drives over one million miles weekly. Safety benchmark: crashes with airbag deployment: 83% fewer than human drivers - Based on Waymo’s first 50 million autonomous miles in its operating domain. Safety benchmark: incidents with injury: 81% fewer than human drivers - Waymo’s reported safety study over the first 50 million miles. Safety benchmark: police-reported incidents: About 64% fewer than human drivers - Waymo’s reported comparison to human driving in similar conditions. Third-party claims reduction: ~80% less property damage or injury claims - External insurance company validation (Swiss Re mentioned). Phoenix territory: Over 300 square miles - Largest operational area, including airport service. San Francisco territory: About 55 square miles - Full city coverage plus some extra area. Los Angeles territory: About 90 square miles - Operational service area. Austin territory: About 37 square miles - Relatively recent expansion area. Waymo research leadership tenure: Since summer 2018 - Drago has led the Waymo research team since 2018. Waymo history with the host: 4 years and 1 month since last appearance - The conversation referenced their prior interview in February 2021.

Pivotal Quotes: "The mission at Waymo is be the world's most trusted driver." — Dragomir Anguilov: Explaining why safety underpins every research and product decision. "There is some unique properties of the problems we're solving... you need to understand the world in 3D even better... very long memory over a scene... think how to prevent hallucinations." — Dragomir Anguilov: Describing why autonomy needs specialized foundation-model adaptations. "We want to be at the forefront and take the benefits." — Dragomir Anguilov: On Waymo’s approach to absorbing foundation-model advances while adapting them to driving.

Implications: Autonomous driving is moving toward larger foundation models plus rigorous simulation and safety engineering. The winners will likely blend world knowledge, 3D/multimodal reasoning, and strong validation—not rely on pure end-to-end black boxes.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast