Episode Summary
Executive Summary: Eric Bernhardsen traced Modal’s origin from his data/infra background at Spotify and Better to a general-purpose cloud platform optimized for fast, serverless execution of code and AI workloads. The conversation centered on why data and AI teams need better developer experience, how Modal’s custom container/runtime stack enables rapid GPU-backed execution, and where the company fits versus cloud providers, Replicate, and traditional serverless offerings.
Main Topics: Eric’s background and path to Modal (Priority: 5/5): Eric discussed his Swedish roots, programming competitions, KTH thesis work at Spotify on music recommendation, seven years at Spotify, and six years scaling Better as CTO before founding Modal in 2021. Why Modal exists: developer productivity for data/AI teams (Priority: 5/5): He argued that data teams are underserved by generic infrastructure and need tools that improve feedback loops, simplify environments, and unify app code with infrastructure code. Custom container/runtime infrastructure (Priority: 5/5): Modal’s core technical advantage comes from building its own container file system, scheduler, and fast startup path so code can run in the cloud within seconds, not minutes. AI inference, fine-tuning, and bursty workloads (Priority: 5/5): The discussion emphasized that serverless fits AI inference especially well, while Modal also supports bursty fine-tuning, batch embeddings, web scraping, video processing, and other parallel workloads. Platform strategy and competition (Priority: 4/5): Eric positioned Modal as a second-layer cloud provider on top of AWS/GCP/Oracle, differentiated from Replicate by serving custom workflows and large-scale production use cases rather than off-the-shelf model APIs. Sandbox and platform-for-platforms direction (Priority: 4/5): Modal’s Sandbox feature began as a safe code-execution tool for LLM apps and is evolving into a backend other companies can programmatically build on, effectively 'functions as a service as a service.' Company building, talent, and long-term cloud economics (Priority: 4/5): Eric reflected on hiring competitive programmers, the importance of talent density, why Modal avoids owning hardware, and how cloud/AI price wars may benefit consumers even if margins are initially negative.
Key Arguments: Data and AI teams need their own infrastructure primitives because Kubernetes and generic cloud tooling are optimized for backend services, not bursty, GPU-heavy, environment-sensitive workloads. Developer productivity is best measured by feedback loops; Modal aims to shrink the loop from code write to cloud execution to seconds. Modal’s technical moat comes from deep infrastructure work: custom file system, scheduler, container startup optimizations, checkpointing, and cloud orchestration. Serverless is a strong fit for AI inference because prompts in and outputs out create a favorable IO pattern, especially when paired with GPUs. Modal’s value is highest for custom models, custom code, and complex workflows, not commodity open-source model endpoints where switching costs are near zero. The company sees itself as a cloud provider one layer above AWS/GCP/Oracle, similar in spirit to Snowflake’s position above the raw public clouds. The Sandbox product is expanding Modal from a developer-facing runtime into a platform that other platforms can embed as execution infrastructure. Owning hardware is unattractive for a venture-backed startup because it reduces capital efficiency and complicates depreciation, operations, and scaling. Competitive programming and similar deep practice-oriented skills correlate with success at Modal because the company’s problems are algorithmically and operationally hard.
Data Points: Spotify team size at thesis time: ~30 people - Eric described Spotify as an obscure startup when he joined for his master’s thesis. Spotify tenure: 7 years - He spent seven years at Spotify building data tooling, recommendation systems, and infrastructure. Better tenure: 6 years - He served as CTO of Better and helped scale the company from a small team to thousands of employees. Better growth: 10 people to 10,000 - Eric described Better’s rapid expansion during his tenure. Better current size after contraction: ~1,000 - He noted the company later shrank significantly after the pandemic-era collapse. Modal founding year: 2021 - Eric said he started Modal in 2021 after leaving Better. Seed round: 2022 - He referenced a seed round with Sarah from Amplify. Series A: announced last year - He said Modal announced a Series A with Redpoint. Container startup latency: ~1–2 seconds - Modal’s early container/file-system work enabled cloud execution within seconds. GPU fan-out: 100 GPUs in a few seconds - Eric said Modal can fan out to at least 100 GPUs quickly. Large-scale fan-out: thousands of GPUs - He said Modal can fan out to thousands of GPUs given a couple of minutes. Large production scale: many thousands of GPUs - He noted Modal has run many thousands of GPUs for backfills and large customer workloads. Slack bot fine-tuning demo: a few hundred lines of code - EricBot was described as a compact end-to-end Modal app. Batch embeddings demo: all of Wikipedia in 15 minutes - He mentioned a planned blog post showing large-scale parallel embeddings on Modal. Competitive programming medal: IOI gold medal - Eric mentioned receiving an IOI gold medal about 20 years ago. Developer productivity claim: ~10x per decade for four decades - He argued developer productivity has increased exponentially over time. Cloud clouds supported: AWS, GCP, Oracle - Modal runs on major public clouds and currently uses Oracle as well. GPU utilization issue: 30–50% typical utilization - The hosts referenced common underutilization in GPU deployments; Eric discussed Modal’s higher utilization approach.
Pivotal Quotes: "I just want to solve that problem. Like, I want to, you know, build something that lets you run things in the cloud and, like, retain the sort of, you know, the joy of productivity as when you're running things locally." — Eric Bernhardsen: Explaining the original motivation for Modal and its focus on fast cloud execution. "Modal is kind of a second layer of cloud provider... we're building a cloud provider. Like, you know, we're like a multi-tenant environment that runs the user code." — Eric Bernhardsen: Describing Modal’s positioning above public cloud infrastructure. "The beauty of Modal is, like, you can run almost anything on Modal." — Eric Bernhardsen: Discussing the platform’s general-purpose compute model and broad workload support.
Implications: Modal is betting that AI and data infrastructure will reward deep developer-experience improvements, not just cheaper tokens or generic Kubernetes wrappers. If it succeeds, more teams may build on higher-level cloud platforms that abstract away container and GPU complexity.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast