The a16z Podcast
The a16z Podcast

World Models, Robotics, and the Future of 3D AI

World Labs co-founder Justin Johnson joins MTS hosts Theo Jaffee and Sophia Puccini to discuss Atlas, World Labs’ latest world model, and the broader case for AI systems that understand and interact with the physical world. Justin explains how Atlas approaches three core tasks: generating new worlds

Featured Speakers

a16z HostJustin Johnson Guest

Topics Discussed

Episode Summary

Executive Summary: Justin Johnson explains World Labs’ Atlas as a foundational world model for generating, reconstructing, and simulating physical spaces from text, images, and sparse captures. The discussion contrasts Atlas with video models and prior Gaussian-splat systems, emphasizing precise camera control, 2D/3D flexibility, and applications in gaming, VFX, VR, construction, and robotics.

Main Topics: World models as a new model class (Priority: 5/5): Johnson frames world models as the spatial/physical counterpart to language models: general horizontal engines for understanding and simulating the real world. Atlas capabilities: generation, reconstruction, simulation (Priority: 5/5): Atlas can generate novel environments, reconstruct real spaces from sparse photos, and simulate scenes for tasks like robotics. Spatial control and long-horizon consistency (Priority: 5/5): The model uses grounded 3D reference images and precise camera steering to reduce drift, hallucination, and loss of structure over longer outputs. 2D vs 3D outputs and the role of Gaussian splats (Priority: 4/5): World Labs argues Atlas directly generates pixels when needed, while explicit 3D representations remain useful for workflows, client rendering, and interoperability. Applications across creative industries and robotics (Priority: 5/5): The conversation covers VFX, games, VR/AR, architecture, construction, and robot training/evaluation via real-to-sim pipelines. Product philosophy and human creative control (Priority: 4/5): World Labs wants tools that give directors and creators fine-grained control rather than slot-machine-style generation. Future of consumer and agentic use (Priority: 3/5): Atlas is positioned as an API-friendly model that can be consumed by humans and agents, potentially enabling fast iteration on interactive experiences.

Key Arguments: World models are to physical space what language models are to text: a general-purpose engine for understanding and manipulating a different modality. Atlas is not merely a product; it is the base model that will power future World Labs applications. Atlas’ three core functions are generation, reconstruction, and simulation, each serving distinct use cases. Grounding reference images in 3D space and giving exact camera control improves long-range consistency and reduces generative drift. Explicit 3D representations like Gaussian splats are still valuable, especially for real-time rendering, mobile/VR, and existing production pipelines. For robotics, reconstructing an actual environment can improve last-mile evaluation and fine-tuning in the exact deployment setting. World models may enable new creative workflows where users and agents collaboratively build games, scenes, and immersive experiences. Creative work will remain human-led, but AI will become a powerful new tool in the same way animation and CGI transformed film.

Data Points: Input photos for reconstruction: as few as 1 photo; up to 100+ photos - Atlas reconstruction can model an existing real-world space from sparse imagery, with fidelity improving as more photos are provided. Video length example: up to a minute - Johnson describes examples where Atlas maintains control over long video generations by leaving breadcrumb reference images in 3D. Launch timing: today - Atlas was announced during the conversation as the world’s first multimodal world model from World Labs. Tripods/cameras for bullet time: 3 iPhones - Atlas can create bullet-time style freeze-and-fly-around shots from synchronized capture with only three phones. Robotics setup time vision: 5 minutes - Johnson imagines a future workflow where a user could take a few photos, upload them, and quickly generate a simulation for robotics training. Foundation model adaptation: 2 pathways - They discuss two tech trees: explicit 3D reconstruction versus direct frame generation without explicit 3D in the loop.

Pivotal Quotes: "Language models are these general horizontal engines for processing streams of discrete text, discrete tokens." — Justin Johnson: He uses this analogy to define world models as the equivalent general engine for physical/spatial understanding. "What Atlas can do is sparse reconstruction. I can take just a few photos of a space, as few as one, up to 100 or more." — Justin Johnson: He explains one of Atlas’ core capabilities: reconstructing real spaces from sparse visual input. "We never wanted to be building these generative slot machines." — Justin Johnson: He describes World Labs’ emphasis on creative control and directability rather than random generation.

Implications: Atlas points to a future where creators, developers, and robots work from editable spatial models instead of fixed assets. If it scales, it could reshape VFX, games, VR, and robot training through fast, controllable world generation.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast