No Priors
No Priors

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

Baseten CEO and co-founder Tuhin Srivastava sits down with Sarah Guo and Elad Gil to discuss the rapid growth of AI inference demand, Baseten’s 30x growth, and why inference is becoming the strategic “last market.” Tuhin Srivastava argues the application layer will persist because companies with uni

Featured Speakers

Tuhin Srivastava Guest

Topics Discussed

Episode Summary

Executive Summary: The conversation centers on Base 10’s rapid growth in AI inference as demand shifts from vanilla model hosting toward custom, post-trained, enterprise-specific systems. Tuhin Srivastava argues inference is the “last market,” driven by open-source model quality, specialized workflows, and Jevons-style demand expansion. He also details severe compute scarcity, the operational complexity of running inference at scale, and why software, supply access, and post-training expertise are now strategic advantages.

Main Topics: AI inference market expansion and Base 10’s growth (Priority: 5/5): Srivastava explains that Base 10 has scaled 30x in a year because AI adoption is broadening from experimentation to production, with more customers using AI everywhere and needing inference infrastructure. Why the application layer will persist (Priority: 5/5): He argues application companies will survive because they control unique user/workflow signals that frontier labs cannot easily access, enabling differentiated post-training and long-horizon agents. Custom models, post-training, and enterprise workflows (Priority: 5/5): The discussion emphasizes that most real usage is not vanilla open-source weights but customized models tuned for quality, latency, or domain workflows, especially in high-scale AI-native companies. Compute scarcity and capacity constraints (Priority: 5/5): Srivastava describes a severe supply crunch across GPUs and clouds, mid-90s utilization, long contract requirements, and operational diligence around providers that can actually run reliable inference workloads. Open-source, Chinese models, and geopolitical competition (Priority: 4/5): He says frontier open-source models from Chinese labs are strong and economically valuable, while also making the case that the U.S. should develop its own open-source models rather than ignore the competition. Multi-chip future and NVIDIA’s current advantage (Priority: 4/5): He expects diversification over time, including inference-specific chips, but says NVIDIA remains dominant in the near term because of CUDA, supply chain execution, and ecosystem maturity. Operating as an inference cloud and scaling culture (Priority: 4/5): Base 10’s product thesis is to build the inference-to-post-training loop, partner where needed, and maintain an incident-driven, high-accountability operating culture suited to always-on infrastructure.

Key Arguments: AI inference demand is exploding because open-source models are now good enough and customers increasingly want to own specialized inference. The application layer remains defensible when companies possess unique user data and workflow signal that labs cannot replicate. Most enterprise-relevant AI usage is custom: customers modify models for quality, performance, latency, and domain-specific behavior. Inference and post-training are converging into a single flywheel: inference produces data, evals, and reward signals that improve future models. Compute scarcity is real and underappreciated; winning inference businesses need supply access, capital efficiency, and operational excellence. Open-source Chinese models are strategically important and economically beneficial to U.S. users, but the U.S. should still build its own open-source stack. Lower inference costs do not shrink demand; they increase it by enabling longer-running agents and more intelligence in products. NVIDIA remains the most practical near-term platform because its ecosystem and deployment velocity are hard to beat.

Data Points: Base 10 growth: 30x over the last year - Describes the company’s scale-up in the AI inference market. Expected revenue: More than $1 billion this year - Referenced as an expectation for Base 10’s current year revenue. Inference share of tokens: 95%+ - Nearly all tokens served on Base 10 are on dedicated/custom inference business. Business lines: 3 businesses - Dedicated inference, shared inference, and training. Cloud footprint: 18 different clouds - Base 10 runs compute across multiple clouds and neoclouds. Cluster count: 90 clusters worldwide - Distributed infrastructure used to satisfy reliability and capacity needs. Typical utilization: Mid-90s - Running clusters at very high utilization due to severe supply constraints. Customer retention: Top 30 customers have never churned - Used to illustrate stickiness of inference plus software. Net dollar retention: 400% annual NDR - Cited as evidence of strong expansion and stickiness. Contract terms for 1,024 B200s: 3-5 year contract with 20-30% TCV prepay - Illustrates how hard it is to secure large GPU capacity now. DeepSeek cost comparison: ~20% of OpenAI/Anthropic production cost - Srivastava says comparable performance/latency may be available at much lower cost.

Pivotal Quotes: "The application layer exists for a number of reasons." — Tuhin Srivastava: Explaining why AI-native application companies can remain defensible against frontier labs. "If you think about the economics here... you could run DeepSeek probably 20% of the cost of running OpenAI or Anthropic models in production." — Tuhin Srivastava: Arguing that cheaper frontier-quality intelligence creates major market expansion if accessible. "In a world of constrained compute, the number one thing to own is compute." — Tuhin Srivastava: Summarizing the strategic importance of supply access in AI infrastructure.

Implications: The AI stack is shifting from generic model access to specialized, continuously improving inference systems. Winners will combine compute access, post-training expertise, and workflow data. For customers, cheaper intelligence means more AI everywhere; for infrastructure vendors, capacity and operations are now core strategy.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors