Episode Summary
Executive Summary: The conversation traces Modal’s evolution from a serverless runtime and developer-experience company into a broader AI infrastructure platform centered on GPU inference, sandboxed agent execution, and elastic training. Akshat explains that Modal’s core insight is colocating code, infrastructure, observability, and networking so AI workloads can scale up and down seamlessly without Kubernetes complexity. The discussion highlights speculative decoding, auto-endpoints, private networking, multi-node training, and the company’s strategy of enabling agent-native primitives across the model lifecycle.
Main Topics: Modal’s origin: from runtime to AI infrastructure (Priority: 5/5): Akshat recounts how Modal began as a better runtime to solve Kubernetes pain, especially for bursty workloads, custom images, and poor developer experience. GPUs were added before ChatGPT, initially for classical inference use cases. Developer experience evolves into agent experience (Priority: 5/5): The team has shifted from optimizing for humans writing YAML and code to optimizing for agents that need to modify decorators, inspect logs, and operate infrastructure autonomously. Modal now treats agent experience as a first-class design goal. Elastic inference as core product-market fit (Priority: 5/5): Modal’s strongest traction is in elastic inference for custom models and open-source model deployment, especially for companies with unpredictable regional and diurnal traffic patterns. GPU snapshotting and autoscaling are key differentiators. Sandboxs and agent primitives (Priority: 4/5): The company built sandboxes early, before the market fully recognized agent workflows, and now offers features like sidecars, file systems, networking controls, and isolated execution for production agents. Inference optimization and Auto Endpoints (Priority: 5/5): Modal is open sourcing inference improvements such as DFlash speculative decoding and launching Auto Endpoints to make frontier-grade performance accessible without requiring users to manually tune inference stacks. Networking, multi-node training, and distributed systems (Priority: 4/5): Modal has built private networking, RDMA support, and overlay networking to enable multi-node training and tightly coupled workloads, while also supporting use cases like egress control, proxies, and co-located services. Capacity strategy and pricing tiers (Priority: 4/5): The discussion covers compute planning, fungible GPU capacity across regions/providers, reliability layers, and future batch pricing for customers that can tolerate latency in exchange for lower cost.
Key Arguments: Kubernetes is a poor fit for bursty, specialized AI workloads; Modal’s runtime is designed to solve that fundamental mismatch. A better developer experience came from colocating infrastructure configuration with code through decorators, reducing YAML and making infra more expressive. The same principles that made Modal good for developers now make it good for agents, since agents benefit from code-proximate infrastructure and live feedback. Elastic inference is the company’s strongest current use case because many AI workloads have unpredictable yet patterned traffic and require rapid regional autoscaling. GPU snapshotting and speculative decoding can produce more meaningful inference gains than incremental kernel-level optimization. Modal’s open-source contributions and collaboration with SG Lang aim to upstream improvements rather than lock users in. Production agents need hard security boundaries, observability, and control over networking/storage—not just a higher-level managed agent wrapper. The future of AI infrastructure will include more post-training, multimodal/video workflows, computational biology, robotics, and batch processing—not just LLM APIs.
Data Points: Cloud providers in Modal capacity pool: 17 - Akshat says Modal spans 17 cloud providers / capacity sources globally. GPU snapshotting speedup: “way faster” startup via state snapshotting - Used to describe saving Torch compiler/model state to reduce inference startup latency. Speculative decoding speedup: 2 to 4x - Akshat explains that higher accept length can yield multiplicative speedups over standard token-by-token decoding. Sandbox build timing: May 2023 - Modal built sandboxes in May 2023 before broad agent/sandbox demand emerged. Training cluster networking bandwidth: 3 terabit per second - Akshat cites internal networking capacity needed for multi-node training. Example scale of sandbox bursts: 100,000 sandboxes - He mentions RL rollouts can require extremely large sandbox bursts at times. Batch pricing latency target: next 24 hours - Modal is exploring cheaper batch tiers for customers who do not need immediate results.
Pivotal Quotes: "Modal is a cloud platform that's built for where we've built the primitives from scratch for AI applications." — Akshat: Definition of Modal’s current positioning as an AI-native infra platform. "Why would you have an agent read through hundreds of Kubernetes files and like write YAML that's not even typed when it can basically make a couple of changes in a decorator?" — Akshat: Explaining why Modal’s code-first abstraction fits agent workflows better than Kubernetes. "The thing that we have actually specialized in is the auto-scaling aspect." — Akshat: Clarifying Modal’s differentiator versus generic inference providers.
Implications: Modal is positioning itself as the infrastructure layer for agentic and multimodal AI workloads. If adoption continues, expect more demand for elastic compute, secure sandboxes, private networking, and managed inference primitives that let agents operate real systems safely and efficiently.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast