Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Earlier this month, world model company Runway introduced GWM Worlds 2, a research preview that “turns high-fidelity video and audio generation into real-time interactive simulation.” Runway calls this an “autoregressive diffusion” model; with autoregressive describing how it generates over time. On

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The conversation traces Runway’s evolution from an artist-friendly tool for early generative models into a frontier lab building video, world, and robotics models. Anastasios argues that scaling video prediction will unlock world understanding, real-time interfaces, and robotics via world models, with controllability, distillation, and multimodal training as key next steps.

Main Topics: Runway’s origin and creative-AI thesis (Priority: 5/5): Anastasios explains how Runway began by exposing difficult open-source generative models to artists, grounded in a belief that creative tools would need to be rethought as generative media improved. From post-production tools to video generation (Priority: 5/5): Runway’s early success came from segmentation/green-screen and image-to-video experiments, then shifted with Gen 1 and Gen 2 toward controllable video generation for filmmakers and creators. Scaling laws and the jump to Gen 3/Gen 4.5 (Priority: 5/5): The team bet that video would follow the same scaling dynamics as language models, investing heavily in training infrastructure and larger diffusion transformers to close the gap after Sora. World models as the next frontier (Priority: 5/5): The discussion emphasizes that video prediction is not just content generation but a way to simulate reality, model physics, and support counterfactual reasoning for real-world tasks. Real-time interfaces and software as pixels (Priority: 4/5): Runway’s interface world model is presented as a future software paradigm where applications are rendered and interacted with directly as generated pixels instead of code-driven UI. Robotics, sim-to-real, and data strategy (Priority: 5/5): Runway frames its robotics work as leveraging video pretraining plus small amounts of robot data to reduce sim-to-real gaps, using world models for synthetic data generation and evaluation. Open source, benchmarks, and industry direction (Priority: 4/5): The transcript covers the Stability AI/Stable Diffusion history, the Cosmos Coalition, benchmark debates, and how artists, enterprises, and AI labs are converging on more controllable tools.

Key Arguments: Runway’s core thesis is that generative models will fundamentally change creative tools, and artists will produce surprising outputs when given flexible generative systems. Video generation should be treated as world simulation, not just content generation; improving video models improves physics understanding and real-world planning. Controllability matters as much as quality: users need camera motion, motion brushes, input frames, and other conditioning signals rather than pure text-to-video. Scaling compute and model size remains the strongest lever; Runway claims predictable gains in physics-related benchmarks as models get bigger. World models can serve both humans and agents: they can power interfaces, educational exploration, and synthetic environments for computer-use and robotics agents. A single pretrained video foundation model can be adapted to robotics with far less robot-specific data than training from scratch. Real-time generation and step distillation are crucial for usability and economics, especially for interactive experiences and front-end replacement. Open multimodal research and better benchmarks are necessary to grow the world-model ecosystem and keep pace with global competition.

Data Points: Runway founding / operating time: ~7–8 years - Anastasios says the company has been around seven years and later calls it “eight years now.” Early generative model timeline: 2016–2017 - He points to early generative models as the inspiration for Runway’s initial thesis. Gallery simulation art project: 2015 - He describes an early interactive art piece that used generated identities and dialogue. Uncanny Road / early image generation: 1K resolution - He says one of the early models generated at 1K resolution using road-scene semantic maps. Runway research scaling bet: 18x100 cluster - He says Runway signed a deal to build a large compute cluster during its Series B era. Gen 1 release: January 2023 - He notes Gen 1 was released in early 2023 as a video-to-video model. Gen 2 release: 2023 - He describes Gen 2 as Runway’s first text-to-video model, built via a two-stage text-to-depth then depth-to-RGB approach. Gen 3 scaling: 10x - He says Runway scaled model size and compute by roughly 10x to build Gen 3 after Sora. Robotics post-training data: hundreds of hours - He contrasts fine-tuning on robot data with pretraining regimes that may use 100,000+ or millions of hours. Character model duration: up to 30 minutes - He says Runway’s character model can generate autoregressive video for up to 30 minutes. Open-ended world model duration: a few minutes - He says the general world model can sustain generation for several minutes before degradation. Video benchmark improvement: predictably improves with compute scale - He says physics IQ scores improve as model scale increases. World model summit date: September 30 - He announces the Runway AI Summit focused on physical AI and real-time video generation.

Pivotal Quotes: "We will need, as a result of those generative models, rethink how creative tools are made." — Anastasios: Describing Runway’s original founding thesis about generative media and creative software. "The human mind is no longer the center of AI, our world is." — Anastasios: Summarizing the shift from language-centric AI to world-centric modeling and simulation. "There is a point where most of content will be generated." — Anastasios: Explaining why Runway believed generative models would become central to media creation.

Implications: Runway sees video AI becoming a general world simulator powering real-time interfaces, robotics, and new software paradigms. If correct, future products will be more interactive, multimodal, and controllable, with video models acting as foundational infrastructure.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast