Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Why Compound AI + Open Source will beat Closed AI

We have a full slate of upcoming events: AI Engineer London, AWS Re:Invent in Las Vegas, and now Latent Space LIVE! at NeurIPS in Vancouver and online. Sign up to join and speak! We are still taking questions for our next big recap episode! Submit questions and messages on Speakpipe here for a chanc

Featured Speakers

Latent.Space HostLin Kiao Guest

Topics Discussed

Episode Summary

Executive Summary: The episode centers on Fireworks AI’s evolution from a PyTorch-first infrastructure idea into a broad compound AI platform for generative AI inference. Lin Kiao explains why inference, multimodal support, and model orchestration matter more than training for most companies, and how Fireworks differentiates through distributed serving, custom kernels, and optimization across quality, latency, and cost.

Main Topics: Fireworks AI origin and two-year company journey (Priority: 5/5): Lin recounts the company’s founding in 2022, the challenges of scaling through the SVB crisis and operational mistakes, and how the team learned to build a company while staying focused on product and go-to-market. PyTorch roots and Meta experience (Priority: 5/5): Lin explains how his Meta background in distributed systems and PyTorch shaped Fireworks’ technical worldview, including the shift from mobile-first to AI-first and the need to support both research and production workloads. Shift from PyTorch cloud to generative AI inference (Priority: 5/5): The company initially imagined a PyTorch SaaS platform, but customer discovery and the ChatGPT moment pushed it toward a verticalized generative AI platform focused on inference rather than training. Compound AI as the product thesis (Priority: 5/5): Lin argues that real applications require multiple models, modalities, and external systems working together, not a single model. This led Fireworks to embrace the compound AI framing and build orchestration around it. Platform breadth: models, modalities, and optimization (Priority: 4/5): Fireworks now offers a wide catalog spanning text, audio, vision, embeddings, image generation, and video, layered on top of Fire Optimizer and a distributed inference engine. Infrastructure differentiation and distributed serving (Priority: 4/5): Lin describes Fireworks’ distributed inference as sharding across GPUs, regions, and hardware types, with custom kernels and load balancing to improve performance and cost efficiency. Industry standards and ecosystem dynamics (Priority: 3/5): The conversation touches on OpenAI-compatible APIs, Meta’s Llama stack, and the role of open source ecosystems in shaping the next layer of AI application infrastructure.

Key Arguments: PyTorch succeeded because it started as a researcher-first tool and then expanded into production, creating a flywheel between open-source adoption and enterprise use. The AI industry is moving from training-centric projects to inference-centric applications because inference scales with users, while training scales with researchers. A horizontal one-size-fits-all platform is insufficient; customers need verticalized and then customized inference setups based on their workload and objectives. Compound AI is necessary because models are probabilistic, specialized, and incomplete; production systems must combine multiple models, APIs, databases, and knowledge sources. Fireworks’ value is not just raw GPU access but a developer-friendly, serverless, optimized layer that abstracts serving complexity and improves quality, latency, and cost. OpenAI-compatible APIs reduce adoption friction and have become a de facto standard, making interoperability a strategic advantage. Open source innovation in audio, vision, and video is accelerating, and Fireworks can build on that ecosystem faster than a closed platform can. Distributed inference should be optimized across GPUs, regions, and hardware types rather than simply splitting a model monolithically across devices.

Data Points: Company age: 2 years - Lin says Fireworks just celebrated its two-year anniversary. Funding rounds: 2 red-hot rounds - The intro references funding from Benchmark and Sequoia Capital. Customer examples: Superhuman, Cursor, Quora, HubSpot - Named as part of Fireworks’ customer list. FireAttention throughput: 15x higher throughput than vLLM - Mentioned as one of Fireworks’ launches advised by SWIX. Public platform launch: August last year - Lin says Fireworks launched its public platform in August. Founding date: September 2022 - Lin says the company started in September 2022. Meta PyTorch transition period: 5 years - Lin says it took about five years to adapt PyTorch for both research and production at Meta. Geographic coverage: North America, EMEA, Asia - Fireworks runs distributed inference across multiple regions. Prize categories for NeurIPS event: 3 - The intro announces three prize categories for the Latent Space Live micro-conference.

Pivotal Quotes: "If we create fireworks and support the industry going through this transition, it will be a huge amount of impact." — Lin Kiao: Explaining the founding motivation for Fireworks AI after seeing the AI-first transition at Meta. "Our goal is to make AI accessible to all app developers and product engineers." — Lin Kiao: Describing why Fireworks focused on easy APIs and inference rather than building a PyTorch cloud first. "In order to really build a compiling application on top of JNI, we need a compound AI system." — Lin Kiao: Summarizing why single-model approaches are insufficient for real-world AI products.

Implications: The episode suggests AI infrastructure is shifting toward multimodal, model-orchestrated, inference-first platforms. For builders, success will depend on abstraction, optimization, and interoperability rather than training from scratch.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast