The a16z Podcast
The a16z Podcast

The Next Frontier of AI Video Is Control

a16z General Partner Jennifer Li sits down with fal co-founder Gorkem Yurtseven and Head of Engineering Batuhan Taskaya to discuss what changes when generative video becomes fast enough to run in real time. They unpack the technical work behind H3 Max, fal’s post-trained version of MiniMax’s open-we

Featured Speakers

a16z Host

Topics Discussed

Episode Summary

Executive Summary: The episode centers on FAL’s H3 Max video model and how post-training plus systems optimization dramatically improved speed, cost, and controllability. The speakers argue that generative video has reached a threshold where real-time, continuous, and professionally directed AI video experiences are now practical for consumers and Hollywood alike.

Main Topics: H3 Max as a breakthrough generative video model (Priority: 5/5): The hosts discuss why H3 Max stood out immediately: it combined frontier-quality video generation with unusually large gains in speed and efficiency, making it feel unlike other video models on the market. Post-training and systems co-design (Priority: 5/5): A major theme is that FAL’s gains came from combining model post-training with kernel, pipeline, and hardware-level optimization, allowing the team to push beyond standard inference efficiency limits. Real-time and continuous video experiences (Priority: 5/5): The conversation highlights new formats enabled by low latency: continuous streaming video, memory across scenes, interactive direction during playback, and long-running sessions with scene coherence. Controllability for creators and Hollywood (Priority: 5/5): The speakers stress that the next frontier is not just speed but control—camera angles, lighting, motion, lip-sync, references, and style controls—especially for professional workflows and studios. Open source ecosystem and rapid experimentation (Priority: 4/5): Because H3 Max builds on an open-weight base model, the ecosystem quickly produced viral demos, livestreams, LoRAs, and custom applications, accelerating adoption and innovation. Market demand and compute economics (Priority: 4/5): They frame generative media as a 'token market fit' category with huge demand but compute constraints, where efficiency improvements directly expand usage and create room for more models and workloads. Hollywood adoption and infrastructure readiness (Priority: 4/5): The episode closes on Hollywood’s growing use of AI video, including studio workflows, IP handling, US hosting, and the shift from curiosity to active production adoption.

Key Arguments: Generative media is one of the clearest examples of 'token market fit' because individuals and teams can productively spend large amounts of compute on creative output. H3 Max became a standout model because it was the first truly capable open-weight, next-generation video model that FAL could post-train and optimize. Most of the speedup came from a combination of post-training for fewer denoising steps and systems optimization across the full pipeline, not from a single trick. The team moved inference utilization from roughly 30–40% toward 70–80% by optimizing kernels, model execution, and pipeline components. Video generation is a multi-stage pipeline, including prompt expansion, diffusion/latent generation, decoding, and sometimes upscaling, so every stage offered optimization opportunities. Real-time video unlocks entirely new use cases, including live, continuous, interactive storytelling and a model that can respond while the video is still playing. Continuous memory across scenes is a major breakthrough; the system can retain recent video context and evolve a longer-term representation for coherence. The biggest remaining gap is controllability, not raw speed: professionals want precise control over camera, lighting, motion, and character behavior. Open source accelerates experimentation because creators can build LoRAs, livestreams, style variants, and application layers around the base model. Hollywood adoption is being driven by practical point solutions—extending scenes, changing angles, editing lighting, and reusing IP—rather than full from-scratch generation. Legal and data residency issues matter alongside model quality, and FAL says it has made progress by supporting hosted models and customer IP workflows.

Data Points: Turbo generation speed: 5-second video in 1.5 seconds - H3 Max Turbo public version Relative cost: About 2x less than the standard version - H3 Max Turbo compared with H3 Max Speedup vs original endpoint: ~35x faster - Compared with the original Minimax H3 endpoint Quality retention: 97th percentile of quality - Turbo version described as nearly matching the main model Utilization improvement: 30–40% to 70–80% - Hardware utilization after systems optimization Memory window: Up to 2 minutes of retained scene memory - Continuous video generation with coherence across scenes Continuous output length: Up to 60 minutes - H3 Max Director continuous action-controlled video Conference timing: Next week - Generative Media Conference mentioned as upcoming Market adoption: Most popular video model on the platform, by more than 2x volume - FAL platform usage after launch Human time scale: 16–17 hour work cycles - Describing the internal 24-hour development burst around the launch Model step reduction: From 50 steps to around 20 steps - Example of diffusion-model efficiency optimization Hardware generation uplift: 2x to 3x improvement - Hopper to Blackwell hardware comparison

Pivotal Quotes: "Generative media is, I would say, along with the coding agent market, what we call token market fit." — Gorka Mirdsev: Explaining why video generation is a strong demand category for AI compute "The next month or two is going to be fully focused on what happens when AI video becomes fast enough to generate in real time." — Jennifer Lee: Framing the strategic next step after H3 Max’s speed breakthrough "We have a version called H3 Max Turbo that's public that can generate like a five-second video in like 1.5 seconds." — Gorka Mirdsev: Describing the public faster variant of the model

Implications: AI video is moving from novelty to infrastructure: real-time, interactive, and professionally controllable workflows are becoming feasible. Expect more consumer experiences, more studio adoption, and a bigger premium on control, memory, and reliability over raw generation alone.

🔓 Sign Up for Unlimited Episode Search

About The a16z Podcast

The a16z Podcast discusses tech and culture trends, news, and the future – especially as ‘software eats the world’. It features industry experts, business leaders, and other interesting thinkers and voices from around the world. This podcast is produced by Andreessen Horowitz (aka “a16z”), a Silicon Valley-based venture capital firm. Multiple episodes are released every week; visit a16z.com for more details and to sign up for our newsletters and other content as well!

View all episodes from The a16z Podcast