Episode Summary
Executive Summary: Latent Space’s February 2024 recap argues that the most promising LLM progress is coming from five research directions: long inference, synthetic data, alternative architectures, mixture-of-experts, and online LLMs. The episode also covers Gemini 1.5, Sora, and the Gemini alignment controversy, framing them as signs that context length, multimodality, and model behavior are becoming central product and research battlegrounds.
Main Topics: Long inference and planning (Priority: 5/5): The hosts argue that the next major scaling frontier is not just training bigger models, but spending more inference time on reasoning, search, planning, and iterative refinement. They cite chain-of-thought, flow engineering, SGLang, and AlphaGeometry as examples of this direction. Synthetic data as a scaling lever (Priority: 5/5): They discuss synthetic data as a way for models to improve themselves, especially in domains with verifiable outputs like math and code. They contrast self-improving synthetic data with older distillation-style data generation and note concerns about collapse and loss of diversity. Alternative architectures and long-context models (Priority: 4/5): The conversation covers Mamba, RWKV, Striped Hyena, diffusion transformers, and retrieval-augmented language models. The hosts are skeptical but open-minded, emphasizing that Gemini 1.5’s long-context performance is a major challenge to assumptions about quadratic attention limits. Mixture of experts and model merging (Priority: 4/5): DeepSeek’s MoE design, with many smaller experts and always-on experts, is highlighted as a notable innovation. The hosts also discuss model merging as a cheap, GPU-light way to combine specialized models into stronger systems. Online LLMs and retrieval (Priority: 4/5): They frame online LLMs as a table-stakes feature for modern assistants, especially for recent information, search, and RAG. Gemini’s search-augmented performance and Perplexity’s weaker benchmark showing are used to argue that online access helps but is not a universal boost. Sora, multimodality, and world models (Priority: 5/5): Sora is treated as a major cultural and technical event that forces the field to take video generation and embodied understanding seriously. The hosts debate whether it truly learns world models or mainly produces convincing synthetic video, while acknowledging its potential for synthetic data and future understanding systems. Gemini alignment and prompt rewriting controversy (Priority: 5/5): The hosts discuss the backlash over Gemini’s image outputs and the revelation that prompts were rewritten for diversity. They argue this is a prompt-layer decision, not necessarily a model-level failure, and that it raises broader questions about alignment, transparency, and soft power.
Key Arguments: Inference time is an underused scaling dimension; models can be made better by spending minutes, hours, or longer on reasoning and planning rather than only milliseconds. Synthetic data is promising because models can generate useful training data for themselves, especially where correctness can be verified, but the ceiling and risk of mode collapse remain unclear. Alternative architectures may help with long context and efficiency, but the practical gains are uncertain because attention is not always the dominant compute bottleneck. Mixture-of-experts can improve performance and efficiency, especially when experts are specialized and some are always-on for common knowledge. Online LLMs are useful for freshness and search, but they are not a guaranteed benchmark win and may matter most for news, coding, and rapidly changing domains. Sora matters because it makes video generation feel like a step toward richer world understanding, even if current outputs still contain obvious physical inconsistencies. The Gemini controversy shows that prompt-level policy choices can dramatically shape model behavior and public perception, and transparency about those choices matters. Model behavior can become a form of soft power, because subtle biases in outputs may influence users’ reasoning and beliefs at scale.
Data Points: Long inference report generation time: about 10 minutes - The hosts mention a product example where reports are generated much more slowly than real time as part of long inference. GPT-3 parameter scale: 175 billion parameters - Used as the reference point for the era when scaling model size was the main strategy. Original GPT-3 training ratio: 300 billion tokens for 175 billion parameters - Referenced as an early scaling-law estimate. Chinchilla-style compute-optimal ratio: 20 tokens per parameter - Cited as DeepMind’s compute-optimal training ratio. Modern LLaMA-style data scale: 200x to 2,000x tokens-to-parameters - Used to show how training has shifted toward much more data per parameter. Reddit data deal: $60 million per year - Mentioned as an example of the high value of user-generated data for model training. Gemini 1.5 context window: 1 million tokens - Highlighted as a major long-context capability, including multimodal inputs. Gemini vs. OpenAI context comparison: at least 10x longer - The hosts say Gemini’s context window is at least ten times longer than OpenAI’s offerings at the time. DeepSeek MoE experts: 64 experts - Used to contrast with the more common 8-expert MoE setups. DeepSeek always-on experts: 1 to 2 always-on experts - Described as a design for common knowledge and memory retention. RWKV concurrency claim: 256 to 1,000 concurrent users on a single GPU - Presented as evidence for linear-scaling, attention-free architectures. Typical transformer concurrency: 8 to 16 concurrent requests per GPU - Used as the comparison point for RWKV’s throughput claim. Latent Space audience milestone: 750,000+ downloads - Mentioned during the anniversary segment. Latent Space readership milestone: 1 million unique readers on Substack - Cited as part of the show’s first-anniversary celebration. Julius user base: 300,000 users - Rahul says Julius has grown to this many users in about six months.
Pivotal Quotes: "the remaining low-hanging fruit is inference time" — Alessio: Explaining why long inference and planning are the next major scaling frontier. "we solve hallucination. I'm like no, you reduce it" — Alessio: Critiquing claims that retrieval or contextual AI can fully eliminate hallucinations. "art for no one by no one" — Grimes (quoted in discussion): Used to describe the Gemini image-generation controversy as accidental conceptual art.
Implications: For builders, the message is to optimize for reasoning time, data quality, and product-specific performance rather than chasing raw benchmark wins. For the industry, long context, multimodality, transparency, and alignment policy are becoming core competitive and cultural issues.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast