Episode Summary
Executive Summary: Eden Ha describes his path from computer vision and NVIDIA’s Cosmos to xAI, explaining how fast model iteration, strong infra, and small teams enabled rapid video, image, audio, and world-model releases. The conversation centers on how video generation depends heavily on language-model prompting, why storage/compute are major bottlenecks, and why the next frontier is real-time, interactive, long-horizon “video agents” and eventually world models.
Main Topics: Career path and transition to xAI (Priority: 5/5): Eden explains moving from Cosmos at NVIDIA to xAI, motivated by scaling laws, compute needs, and the chance to build video/multimodal systems from scratch. How frontier video models are built (Priority: 5/5): Discussion of the pipeline: human and synthetic captioning, image-model bootstrapping, VAEs/tokenizers, diffusion/flow-matching training, and why video models often start with image models. Infrastructure, iteration speed, and bottlenecks (Priority: 5/5): He emphasizes that strong infra, low communication overhead, and many training iterations per day matter more than new algorithms, while compute and storage become the limiting factors. Real-time video generation and world models (Priority: 5/5): The conversation defines world models as interactive, real-time, long-horizon video systems with examples like Flipbook and Neural OS, and explores the path toward generative UI and interactive environments. Audio-video joint generation and multimodal alignment (Priority: 4/5): Eden discusses the challenges of aligning audio, video, and text, including time alignment, discrete vs continuous audio, and the difficulty of captioning music precisely. Video agents vs. pure model scaling (Priority: 5/5): A major argument is that much of the quality gain now comes from language-model reasoning, prompt rewriting, tool use, and post-production orchestration rather than only better video diffusion models. Context management, distillation, and efficiency (Priority: 4/5): They compare long-context handling in video and language models, including reference video, frame packing, step distillation, consistency models, and the role of pruning/compression.
Key Arguments: The biggest performance gains in modern video generation often come from language-model prompt rewriting and agentic orchestration, not just from improving the diffusion/video backbone. Video model quality is constrained by data quality and caption alignment; detailed human/synthetic captions are necessary because raw internet video metadata is weakly aligned to visual content. Training video models is expensive in both compute and storage; data pipelines can become I/O-bound and storage/egress costs can reach millions monthly at scale. A small, highly aligned team with minimal communication overhead can move much faster than a larger group, especially when iteration speed is high. World models should be defined as real-time, interactive, long-horizon video systems rather than just “video generators.” Temporal compression and context selection are essential because naive full-history conditioning explodes token length; selective reference mechanisms are a practical bridge. Video agents can use deterministic tools like FFmpeg, Photoshop, and editing workflows to produce production-grade output better than end-to-end generation alone. The next major shift may be models that manage their own context and harnesses, making pruning, memory, and self-modification part of the model rather than external heuristics.
Data Points: Timeline for xAI video team launch: ~3 months - Eden says xAI built the initial video/multimodal stack and released the first model in about three months. Move from NVIDIA to xAI: Mid-2025 - He says he joined xAI around mid-2025 after working on Cosmos at NVIDIA. Cosmos 1 timing: End of 2024 - He places Cosmos 1 around late 2024. Video token count for 5 seconds: 50K–60K tokens - Used to illustrate why long video contexts quickly become infeasible. Reference video inputs: Up to 7 images - He describes xAI’s reference-to-video feature as allowing up to seven conditioning images. Operating cost example: H100 at $1/hour - Used as a rough illustration of expensive inference/training compute. Monthly GPU cost example: $240/month - Based on one H100 used 8 hours/day for 30 days at $1/hour. AWS S3 storage cost example: ~$100K/month per petabyte scale noted - Part of the discussion on how expensive it is to store large video corpora and features. AWS egress at 5 PB: ~$230K - Referenced as a rough cost for downloading 5 petabytes of data. Video model size example: 19B parameters - He cites LTX as an example of a dense video model at this scale. MOE scale example: 20B active / 100B+ total - He notes that some video-model approaches explore MoE designs similar to medium-scale LLMs. Training data scale: Tens of trillions of visual tokens - He says Cosmos disclosed training on tens of trillions of visual tokens. Inference steps for production: 4–8 steps - He says Cosmos-style models can be distilled to run in a few steps for production. Real-time voice interaction target: ~200 ms - He uses this as a rough latency target for interactive digital-human style systems. CSGO response example: Sub-10 ms to a few ms - He uses gaming latency to illustrate real-time constraints for interactive world models. Compute improvement claim: 100x–1,000x every 12–18 months - A claim made in the discussion about frontier language-model performance/compute efficiency trends.
Pivotal Quotes: "The most important thing is the talent. Everyone was very strong and clever, very close with each other towards a common goal." — Eden Ha: Explaining why xAI could ship a video model so quickly with a small team. "World model is like real-time interactive long-horizon videos." — Eden Ha: His personal definition of a world model and the basis for his product/technical roadmap. "The gain comes from language model, not coming from the video model itself." — Eden Ha: His core thesis that prompt rewriting, tool use, and reasoning drive much of the current improvement in generative media.
Implications: The conversation suggests the frontier is shifting from standalone generation to agentic, tool-using systems with memory, context control, and real-time interaction. For builders, language intelligence, infra, and data pipelines may matter more than raw diffusion advances.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast