No Priors
No Priors

Speed will win the AI computing battle with Tuhin Srivastava from Baseten

At a time when users are being asked to wait unthinkable seconds for AI products to generate art and answers, speed is what will win the battle heating up in AI computing. At least according to today’s guest, Tuhin Srivastava, the CEO and co-founder of Baseten which gives customers scalable AI infra

Featured Speakers

Tuhin Srivastava Guest

Topics Discussed

Episode Summary

Executive Summary: The episode examines Base 10’s role in AI infrastructure, especially inference, and why speed, reliability, and software optimization are becoming decisive advantages. Tuhin Srivastava argues that AI adoption is accelerating faster than expected, inference is more repeatable and customer-facing than training, and enterprises will increasingly prefer buying infrastructure over building it as workloads scale and costs rise.

Main Topics: Base 10’s product and mission (Priority: 5/5): Base 10 provides fast, scalable AI infrastructure for teams building on large models, with a current focus on inference and a long-term goal of expanding beyond it. Why inference differs from training (Priority: 5/5): Tuhin contrasts inference and training across latency, reliability, networking, workflow repeatability, and deployment needs, arguing inference is more operationally demanding and customer-visible. Performance optimization as the core battleground (Priority: 5/5): The conversation explores how low-level software work, batching, decoding techniques, and close NVIDIA collaboration drive throughput and latency gains in inference. Enterprise adoption and the buy-vs-build shift (Priority: 4/5): The speakers discuss how enterprises are moving from experimentation to deployment, why speed matters most, and why many customers now prefer buying infrastructure instead of building it. Model size, local deployment, and commoditization (Priority: 4/5): They discuss smaller, more efficient models, distillation, and local execution, suggesting some language-model inference will commoditize while optimization remains important. Hardware supply, GPU shortages, and heterogeneity (Priority: 4/5): The discussion covers ongoing access constraints for premium GPUs, long procurement cycles, and skepticism about near-term multi-vendor hardware abstraction beyond NVIDIA. Industry structure, defensibility, and market dynamics (Priority: 3/5): The hosts and guest debate whether AI businesses will become oligopolies, where defensibility comes from, and how workflows, data, and contracts affect competition.

Key Arguments: Inference has more repeatable customer workflows than training, so infrastructure products can standardize around common deployment, versioning, CI/CD, and cold-start issues. Reliability and latency matter more in inference because user-facing applications cannot tolerate failures or slowdowns; training can be more forgiving. AI market demand, not just product quality, determines adoption speed; the 2022–2023 surge validated the space and accelerated customer urgency. Performance gains come from combining new research, low-level kernel work, batching techniques, and better hardware, especially NVIDIA’s latest stack. Smaller, more specialized models will increasingly run locally or in more efficient environments, reducing some cloud dependence and commoditizing parts of the stack. Enterprises will likely underinvest in AI over the next 12–18 months relative to the opportunity, but materially underinvest over 3–5 years. Buy-vs-build is shifting toward buy because infrastructure is repeated, while proprietary value lies in models, data, and workflow; building infra slows teams down. Premium GPU access remains constrained, and hardware heterogeneity is attractive in theory but difficult in practice because software and debugging complexity are still centered on NVIDIA.

Data Points: Base 10 operating history: 4.5 years - Tuhin says the company has spent the last four and a half years building Base 10. Inference benchmark throughput: 90 tokens/sec to 100+ to 200–400 tokens/sec - He describes how state-of-the-art serving performance has improved over the last six months. Customer latency target: sub-300 ms / sub-200 ms responses - Bland AI uses Base 10 to co-locate workloads and serve call-center SDK requests quickly. Enterprise AI spend example: tens of millions of dollars - He cites Pfizer earmarking this scale of investment over 12–18 months. Customer growth example: four engineers to mid-hundreds of thousands of dollars annual spend - A customer with fewer than 1,000 users grew to significant annual inference spend by year-end. Migration speed example: 36 hours - A four-person AI infrastructure team migrated workloads to Base 10 in this time. Token volume example: 1 billion tokens/day - A six-person chatbot company reportedly processes this volume through its system. Time horizon for enterprise adoption: 12–18 months vs. 3–5 years - Tuhin says near-term enterprise adoption may be overestimated, but long-term adoption is underestimated. Common hardware progression: T4s and A10Gs to A100s to H100s - He describes the progression in customer GPU usage as workloads have scaled. Performance stack example: Whisper / Faster Whisper / Whisper on TRT - Used as an example of optimizing speech models for real-time use cases.

Pivotal Quotes: "Speed is actually your number one advantage." — Tuhin Srivastava: He explains why AI teams must move quickly to remain competitive in a fast-changing market. "What is proprietary to them is models, data, and workflow. What is repeated for them is infrastructure." — Tuhin Srivastava: His core argument for why customers should buy infrastructure instead of building it themselves. "After payroll, this is our second biggest expense for the year." — CTO of a customer (as quoted by Tuhin): Illustrates how large inference spending can be for even a relatively small AI company.

Implications: AI infrastructure is becoming a major budget line, not a niche tooling category. Expect faster enterprise adoption, rising inference spend, continued GPU pressure, and more demand for managed, specialized infrastructure as speed and reliability become strategic advantages.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors