Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

The Four Wars of the AI Stack (Dec 2023 Audio Recap)

Note for Latent Space Community members: we have now soft-launched meetups in Singapore, as well as two new virtual paper club/meetups for AI in Action and LLM Paper Club. We’re also running Latent Space: Final Frontiers, our second annual demo day hackathon from last year. Edit from March 2024: We

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: The episode is a year-end/2024 kickoff recap of AI’s major battlegrounds: data rights and synthetic data, GPU-rich vs GPU-poor inference economics, multimodality, and the tooling/RAG stack. The hosts argue that AI progress is increasingly constrained by data quality, compute costs, and productization, while open-source, code, and agents were less central than expected. They also highlight emerging shifts in hardware, context capture, and production deployment.

Main Topics: The 'four wars' framing of AI in 2023 (Priority: 5/5): The hosts organize the AI landscape into four major battlegrounds: data, GPU-rich vs GPU-poor compute, multimodality, and RAG/tooling. They use this framework to explain where money, talent, and strategic competition concentrated during the year. Data rights, attribution, and synthetic data (Priority: 5/5): They discuss the legal and economic fight over training data, including the New York Times lawsuit, API lockups by UGC platforms, and the rise of synthetic data as human data becomes scarcer and more restricted. GPU inference economics and the Mixture-of-Experts shift (Priority: 5/5): The conversation covers the price war in inference, benchmark disputes, and how MoE models like Mixtral change hardware and serving requirements. They argue many providers are likely losing money on aggressive token pricing. Multimodality and consumer AI products (Priority: 4/5): They review image, voice, and hardware modalities, emphasizing Midjourney’s success, ElevenLabs’ growth, and the idea that multimodal AI creates new consumer markets rather than replacing existing ones. RAG, vector databases, and AI tooling (Priority: 4/5): The hosts argue that RAG remains essential, but the market is shifting from simply storing vectors to making them operationally useful. They see databases, frameworks, and ops tools converging and competing for control of the AI stack. Coding agents and the limits of outer-loop automation (Priority: 4/5): They distinguish between inner-loop coding assistance and outer-loop autonomous agents, concluding that current progress is strongest inside the IDE and that fully autonomous non-technical code generation remains premature. Emerging architectures and the 'sour lesson' (Priority: 3/5): They debate state-space models, Mamba, RWKV, and other transformer alternatives. One speaker introduces the 'sour lesson': AI may not need to resemble human cognition to be useful, but the empirical case is still uncertain.

Key Arguments: Training data is becoming a strategic asset, and the industry is moving from open scraping toward licensing, partnerships, and first-party data creation. Synthetic data will be a major theme because human data is increasingly locked up, but its value depends on whether it adds useful signal rather than merely resampling model outputs. Inference pricing below roughly cost-recovery levels is likely unsustainable, so many providers may be subsidizing growth or pursuing a long-term land-grab strategy. MoE models create new serving constraints because all weights must remain available even if only part of the model is active, pushing innovation in kernels, batching, and quantization. RAG is not a passing trend; it is a necessary production pattern for most AI applications, especially where infinite context is impractical or unreliable. Open-source AI is not really a 'war' because nearly everyone wants it to improve; the real conflict is around who controls inference and distribution. Code models matter, but the market is fragmented and harder to adopt than text or vision, so they have not become as dominant a battleground as general reasoning or multimodal AI. Outer-loop autonomous agents are still too early; the most reliable value today comes from inner-loop tools that assist developers within existing workflows. The biggest AI winners may be products that capture unique context over time, not just models with the best raw capabilities. Emerging architectures should be judged less by whether they mimic humans and more by whether they deliver better efficiency or useful behavior in production.

Data Points: Monthly recap length: Substack newsletter was so long it caused formatting issues - The hosts joked that the December recap was unusually large and may have broken Substack due to the number of links. Poolside fundraising: $50 million+ seed - Referenced as a notable code-model company fundraise, with much of the capital reportedly spent on GPUs. Mixtral price drop: ~90% in one week - The release of Mixtral triggered a sharp drop in inference pricing and intensified the price war. Midjourney revenue: $200M-$300M ARR - They cite reports that Midjourney reached at least $200M ARR, possibly higher, with a very small team. Midjourney team size: 15-30 employees - Used to emphasize how unusually capital-efficient the company is. Inference breakeven estimate: $0.50-$0.75 per million tokens - They estimate the lowest sustainable cost for serving Mixtral, implying some providers charging below that are likely losing money. Perplexity Mixtral price: $0.56 per million output tokens - Used as a benchmark for near-cost pricing in the inference market. AnyScale price: $0.50 per million tokens - Cited in the benchmark/pricing discussion as likely below breakeven. Octo AI price: $0.50 per million tokens - Included in the list of providers likely pricing below cost. Abacus AI price: $0.30 per million tokens - Included in the list of providers likely pricing below cost. Deep Infra price: $0.27 per million tokens - Included in the list of providers likely pricing below cost. OpenAI/Google competition: 2 major general-purpose model contenders - They frame Gemini as the main credible alternative to OpenAI at the time of recording. Context window example: 100k context considered acceptable - Used to argue that long-context models are less compelling than efficiency gains for many use cases.

Pivotal Quotes: "the four words of the AI stack" — Alessio: Introduces the episode’s organizing framework for the major AI battlegrounds of 2023. "the year of AI in production" — Swix: Used to describe the ongoing shift from demos and experimentation toward real deployments. "stop trying to model artificial intelligence like human intelligence" — Alessio: Defines the 'sour lesson' argument about why AI architectures may not need to resemble the brain to be effective.

Implications: Listeners should expect 2024 to be shaped by data licensing, synthetic data, inference economics, and production tooling. The biggest opportunities likely sit in context capture, RAG, and efficiency gains, while fully autonomous agents and exotic architectures remain promising but unproven.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast