Episode Summary
Executive Summary: Fei-Fei Li and Justin Johnson, co-founders of World Labs, discuss spatial intelligence as the next frontier beyond language models. They introduce Marble, a generative model that creates interactive 3D worlds from text or images, positioning it as both a practical product and a step toward general world models. The conversation covers the history of AI scaling, the role of academia versus industry, the challenges of imbuing models with physics understanding, and the complementary nature of spatial and linguistic intelligence.
Main Topics: The evolution of deep learning and scaling compute (Priority: 4/5): Deep learning's progress is tied to increasing computational power, from AlexNet's GPUs to modern clusters with millionfold more capability. Spatial intelligence vs. language intelligence (Priority: 5/5): Spatial intelligence is complementary to linguistic intelligence, rooted in embodied 3D understanding, and underappreciated compared to language. Marble as a product and foundational model (Priority: 5/5): Marble generates 3D worlds from multimodal inputs, offers interactive editing, and outputs Gaussian splats for real-time rendering, targeting gaming, VFX, and simulation. Academia vs. industry research dynamics (Priority: 3/5): Academia is under-resourced but should focus on wacky, long-term ideas rather than scaling battles. Open science remains important despite commercial pressure. Physics and emergence in world models (Priority: 4/5): Deep learning fits patterns, not causal laws. True physical understanding may require emergent capabilities at scale, not explicit physics injection. The data structure of 3D world models (Priority: 3/5): Gaussian splats are the current atomic unit for Marble, but future models could use tokens or other representations to incorporate dynamics and interactivity. Historical projects at Stanford (Priority: 2/5): Fei-Fei and Justin recount work on image captioning and dense captioning, noting simultaneous independent discovery and the evolution from CNNs to LSTMs.
Key Arguments: Spatial intelligence is a distinct, complementary modality to language intelligence, requiring different data structures and scaling approaches. Marble balances being a practical product with advancing toward general world models for spatial intelligence. Academia should pursue innovative, long-term ideas even with limited compute, as industry labs dominate large-scale training. Emergent understanding of physics may arise from scaling world models, but it remains uncertain whether latent learning will yield causal models. Transformers are fundamentally set models, not sequence models, making them flexible for 3D tokens beyond 1D sequences. Synthetic data from world models like Marble can address data scarcity in robotics and embodied AI training. Human intelligence combines multiple intelligences (spatial, linguistic, logical, emotional), and AI should similarly integrate them. Gaussian splats enable real-time rendering on mobile devices, but new architectures are needed for dynamics and physics simulation.
Data Points: Compute scaling: 1,000x performance per GPU card - From AlexNet era to present, per-card performance increased roughly a thousandfold. Compute scaling: 1,000,000x total compute - Modern models can train on tens of thousands of GPUs, representing a millionfold increase over AlexNet. Vision evolution duration: 540 million years - Nature took 540 million years to optimize perception and spatial intelligence. Language evolution duration: 500,000 years (at most) - Language development is estimated at 500,000 years at most, much shorter than visual intelligence. Human speech token output: 215,000 tokens per day - Speaking 150 words per minute for 24 hours yields about 215,000 tokens. Marble output format: Gaussian splats - Marble natively outputs Gaussian splats for real-time rendering on mobile and VR devices.
Pivotal Quotes: "I think the whole history of deep learning is in some sense the history of scaling up computing." — Justin Johnson: Discussing the evolution of deep learning and the need for massive compute for world models. "Spatial intelligence is the capability that allows you to reason, understand, move, and interact in space." — Fei-Fei Li: Defining spatial intelligence as distinct from language intelligence. "I think the world is big enough to have different approaches." — Fei-Fei Li: On whether World Labs should focus on embodied AI directly or build virtual worlds first.
Implications: World models like Marble could transform creative industries, robotics simulation, and spatial understanding in AI. The horizontal nature of the technology suggests broad future applications. The emphasis on complementary spatial intelligence may influence AI research directions beyond pure language models.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast