The a16z Podcast
The a16z Podcast

Google DeepMind Lead Researchers on Genie 3 & the Future of World-Building

Genie 3 can generate fully interactive, persistent worlds from just text, in real time. In this episode, Google DeepMind’s Jack Parker-Holder (Research Scientist) and Shlomi Fruchter (Research Director) join Anjney Midha, Marco Mascorro, and Justine Moore of a16z, with host Erik Torenberg, to discus

Featured Speakers

a16z Host

Topics Discussed

Episode Summary

Executive Summary: The episode centers on Genie 3, Google DeepMind’s research preview for generating fully interactive, persistent worlds in real time from text. The guests explain how it combines real-time responsiveness, minute-plus memory, strong prompt adherence, and emergent world understanding, while also discussing applications in gaming, training agents, and robotics, and why world models are a distinct frontier from traditional video generation.

Main Topics: Genie 3 as a real-time interactive world model (Priority: 5/5): The panel frames Genie 3 as a breakthrough in generating environments that users can navigate live, contrasting it with earlier non-interactive video generation and emphasizing the “wow” factor of immediate responsiveness. Persistent memory and consistency across time (Priority: 5/5): A major focus is Genie 3’s special memory, which preserves state when users look away and back, enabling persistent worlds and making interactions feel coherent and believable. Applications across entertainment, agents, and education (Priority: 4/5): Speakers repeatedly stress that the model is a general capability layer that could power games, agent training, reasoning environments, and educational simulations, with many downstream uses still undiscovered. Training approach, scale, and emergent behaviors (Priority: 4/5): The team describes how scaling data, compute, and breadth of training improved realism, physics, and behavior; they also note emergent interactions such as doors, water, terrain, and prompt-following in unusual scenarios. Text control and the move away from image prompting (Priority: 4/5): They explain that direct text-to-world generation significantly improved controllability and alignment versus prior image-prompted approaches, and that leveraging broader internal research helped accelerate progress. Robotics and embodied AI potential (Priority: 5/5): Genie 3 is presented as an environment model that could support learning through simulation, helping close the sim-to-real gap and potentially accelerate robotics by generating diverse, safe, data-rich training scenarios. Productization, access, and future model directions (Priority: 3/5): The guests discuss why Genie 3 and VO3 remain separate projects for now, the lack of a public product timeline, and the likelihood that future progress will come from continued scaling plus new ideas.

Key Arguments: Real-time interactivity is a fundamental leap over prior video generation because it turns passive outputs into navigable worlds. Persistent memory is planned for and crucial: the model can maintain state for about a minute, making worlds feel consistent when users revisit areas. The core value is not one single application but a general capability to create worlds from text, from which many applications can emerge. Scaling data and compute produces emergent world behavior, including improved physics, terrain interactions, and stronger prompt adherence. Direct text-to-world generation is more powerful than image prompting because it improves controllability and starts the model in the right representational space. Genie 3 is best understood as an environment/simulator for agents, not an agent itself, making it especially promising for RL and robotics. Robotics progress is bottlenecked by expensive, unsafe, and constrained real-world data collection; simulated worlds can help bridge that gap. The model is currently a research preview, and the team expects broader access eventually, though no concrete public timeline was given.

Data Points: Memory duration: One minute - The current design of Genie 3 supports minute-plus memory, with the team saying the current limit is one minute for this type of memory. Genie 2 memory: A few seconds - Jack described Genie 2 as having only a few seconds of memory, with examples like a robot near a pyramid remaining in view after looking away and back. Project start: 2022 - Jack said the GE/Genie project originally started in 2022, initially focused on RL environment generation. Go milestone: 2016 - Jack referenced AlphaGo reaching superhuman level in 2016 as a landmark in RL. StarCraft milestone: 3 years later - Jack noted StarCraft came three years after Go as another major RL environment challenge. VO2 timing: December, roughly a week after Genie 2 - Jack compared Genie 2 and VO2 as contemporaneous announcements in a busy period. Research preview horizon: 7 months - Jack said the research effort leading to Genie 3 took seven months before the surprising samples were seen.

Pivotal Quotes: "it was totally planned for, but still incredibly surprising when it worked that well" — Jack Parker Holder: Jack explaining the origin of Genie 3’s persistent memory and why the team was amazed by the result. "the real-time component is really important" — Anjane Midha: Discussion of why interactivity changes the experience from passive video to something magical and compelling. "we designed it to be an environment rather than an agent" — Shlomi Fuchter: Explanation of Genie 3’s role as a simulator for learning and robotics rather than an autonomous decision-maker.

Implications: Genie 3 suggests world models may become a core AI platform for simulation, agents, and embodied learning. If access broadens, developers could build new gaming, robotics, and education experiences on top of interactive generated worlds.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast