Lex Fridman Podcast
Lex Fridman Podcast

#306 – Oriol Vinyals: Deep Learning and Artificial General Intelligence

Oriol Vinyals is the Research Director and Deep Learning Lead at DeepMind. Please support this podcast by checking out our sponsors: – Shopify: https://shopify.com/lex to get 14-day free trial – Weights & Biases: https://lexfridman.com/wnb – Magic Spoon: https://magicspoon.com/lex and use code L

Featured Speakers

Lex Fridman HostAurelia Vinialis Guest

Topics Discussed

Episode Summary

Executive Summary: Aurelia Vinialis and Lex explore the future of general AI through multimodal models, arguing that progress will come from scaling general methods, modular reuse, better benchmarks, and richer interaction. They discuss Gato and Flamingo as early systems that unify language, vision, and action, and emphasize that current models are powerful but still far from true lifelong learning, robust meta-learning, or consciousness.

Main Topics: AI interviewers, engaging conversations, and the Turing test (Priority: 5/5): The discussion opens with whether AI could replace the interviewer or interviewee in creating compelling conversations. Vinialis argues AI can augment human creativity and perhaps optimize for engagement, but fully replacing the human side may remove what makes conversation interesting. Multimodal general agents: Gato and Flamingo (Priority: 5/5): Vinialis explains Gato as a general agent trained on language, images, actions, and robotics trajectories, and Flamingo as a modular vision-language model built by freezing a language backbone and adding visual capabilities. Both are early steps toward general-purpose systems. Tokenization and unifying modalities as sequences (Priority: 5/5): A major technical theme is that text, images, and actions can all be converted into tokens and modeled as sequences. This lets a single transformer learn across modalities, with shared weights discovering connections through data rather than handcrafted rules. Meta-learning, prompting, and interactive teaching (Priority: 4/5): He reframes meta-learning as learning to learn through prompting, few-shot examples, and eventually richer interaction. He predicts future systems will be taught tasks interactively rather than retrained from scratch, moving beyond static prompting to feedback-driven learning. Scale, emergence, and the limits of current transformers (Priority: 5/5): The conversation covers scaling laws, emergent abilities, and the idea that certain capabilities appear only after thresholds of scale and data. Vinialis sees scaling as necessary but not sufficient, with architecture, search, and data design still crucial. Benchmarks, engineering, and the human factor in AI progress (Priority: 4/5): He emphasizes that breakthroughs depend on benchmarking, data engineering, compute infrastructure, and the people behind the work. Good benchmarks steer progress, while engineering details and team composition can strongly affect what gets discovered and when. Consciousness, sentience, and the social implications of AI (Priority: 4/5): Vinialis says current systems are not sentient and do not need consciousness to be highly useful, but public perception and future human-AI relationships will force society to address rights, safety, and legal norms around systems that appear alive.

Key Arguments: AI can help create more engaging conversations, but fully removing the human participant may make the interaction less meaningful. General agents are built by converting many data types into sequences and training one transformer on them with a next-token objective. Gato is 'the beginning' because it unifies modalities and tasks, but it is still small and not as strong as specialized systems. Flamingo shows modularity works: freeze a strong language model and attach a visual module rather than retrain everything from scratch. Meta-learning is shifting from narrow benchmark definitions to interactive teaching through language, images, and eventually richer feedback loops. Current models have limited working memory and do not continue learning after deployment, which blocks true lifetime-like experience. Scaling is a real driver of capability, but it does not eliminate the need for better architectures, data curation, and benchmarks. Emergent behavior appears when tasks require multi-step reasoning; progress can look flat until a threshold is crossed. Benchmarks, compute, and engineering infrastructure are not secondary details; they shape what research becomes possible. Consciousness is not required for useful intelligence, but AI systems that communicate as if sentient may create major social and legal challenges. Human researchers, their tastes, and their persistence strongly shape the order and pace of breakthroughs.

Data Points: Gato parameter count: 1 billion parameters - Vinialis describes Gato as relatively small compared with modern large models. Chinchilla parameter count: 70 billion parameters - Used as the frozen language backbone in Flamingo. Flamingo added parameters: ~10 billion parameters - Additional multimodal components were added on top of Chinchilla. Flamingo total size: ~80 billion parameters - Chinchilla plus added visual modules. Current context window mentioned: ~2,000 words - He uses this as a rough example of limited working memory in dialogue models. Magic Spoon protein per serving: 13 to 14 grams - Sponsor segment, cited while discussing the cereal product. Magic Spoon sugar: 0 grams - Sponsor segment. Magic Spoon carbs: 4.9 grams - Sponsor segment. Magic Spoon calories: 140 calories - Sponsor segment. Weights & Biases users: Over 200,000 machine learning engineers and data scientists - Sponsor segment. Shopify entrepreneurs powered: Over 1.7 million - Sponsor segment.

Pivotal Quotes: "Gatto is not the end, it's the beginning." — Aurelia Vinialis: He explains that Gato is an early multimodal general agent, not the final architecture. "I think we lack maybe benchmarks and the technology to have this lifetime-like experience of memory that keeps building up." — Aurelia Vinialis: He identifies persistent memory and continual learning as major missing pieces in current AI. "I personally think probably not to the degree. Of intelligence that this brain can learn, can be extremely useful, can challenge you, can teach you, conversely, you can teach it to do things." — Aurelia Vinialis: He argues consciousness is not necessary for highly capable AI.

Implications: The next AI leap likely comes from multimodal generalists that learn interactively, reuse prior weights, and are judged by better benchmarks. For industry, the winners will pair scale with modularity, engineering, and safety planning.

🔓 Sign Up for Unlimited Episode Search

About Lex Fridman Podcast

Conversations about science, technology, history, philosophy and the nature of intelligence, consciousness, love, and power. Lex is an AI researcher at MIT and beyond.

View all episodes from Lex Fridman Podcast