The TWIML AI Podcast
The TWIML AI Podcast

How to Engineer AI Inference Systems with Philip Kiely - #766

In this episode, Philip Kiely, head of AI education at Baseten, joins us to unpack the fast-evolving discipline of inference engineering. We explore why inference has become the stickiest and most critical workload in AI, how it blends GPU programming, applied research, and large-scale distributed s

Featured Speakers

Philip Kiley Guest

Topics Discussed

Episode Summary

Executive Summary: Philip Kiley argues that inference is now the most important and fastest-moving part of AI infrastructure, with research-to-production cycles measured in hours. The conversation covers why inference engineering is complex, how companies progress from API-based usage to owned or dedicated deployments, the role of hardware/software specialization, and why deep inference knowledge matters for building faster, cheaper, more reliable AI products.

Main Topics: Why inference is the core AI workload (Priority: 5/5): Kiley explains that inference is the sticky, business-critical layer of AI infrastructure because it directly affects product speed, cost, reliability, and user experience. The complexity of inference engineering (Priority: 5/5): Inference requires expertise across GPU programming, distributed systems, quantization, speculation, caching, hardware orchestration, and end-to-end service design. Research-to-production speed in inference (Priority: 5/5): Inference adopts new techniques unusually fast; new papers or model architectures can be implemented and deployed within hours or days, not months or years. Product maturity and deployment choices (Priority: 4/5): Companies typically move from closed-model APIs to hyperscalers, then to dedicated inference providers or in-house/edge systems as needs for cost, capacity, and control increase. Hardware longevity and specialization (Priority: 4/5): Older GPU generations remain relevant for years, while newer trends point toward disaggregation, specialized chips, and workload-specific hardware-software optimization. AI-assisted engineering and the future of the role (Priority: 4/5): AI will accelerate parts of inference engineering like CUDA kernel writing, but Kiley argues humans will still be needed to own uptime and mission-critical systems. Open source runtimes and tooling ecosystem (Priority: 3/5): The field has concentrated around a few major open-source runtimes, with providers typically building on top of them rather than reinventing inference from scratch.

Key Arguments: Inference is the most important workload in AI because it is the part that directly determines whether an application feels magical or sluggish to users. Inference is hard because it combines GPU-level optimization, ML research, and traditional distributed systems at the same time. The time from inference research to production is often hours, making it one of the fastest adoption cycles in technology. As models get more capable and AI use shifts from chat to agents, the number of inference requests per user action multiplies, increasing the need for optimization. Knowing inference helps teams move from treating model behavior as fixed to understanding a performance spectrum with tunable tradeoffs. Companies usually start with closed-model APIs, then move to hyperscalers for cost/capacity relief, and later to dedicated providers or in-house platforms for greater control. Hardware stays relevant longer than many assume: Hopper and even some Lovelace workloads continue to be useful because software ports and model optimization lag hardware release cycles. AI will automate some low-level tasks, but not the accountability and operational ownership needed for high-stakes inference systems. Specialized inference is becoming more important as workloads diversify across text, vision, speech, embeddings, and agentic tool use. Open-source infrastructure is converging around a small set of runtimes, making ecosystem contribution more practical than building everything from zero.

Data Points: Time for medicine research to reach a pharmacy: decades - Used as a contrast to show how quickly inference innovations can move into production. Time for physics/engineering concepts to be applied: years - Contrast with inference’s much faster research-to-production cycle. Time to implement new inference techniques: hours - Kiley says inference research often moves into production within hours. Time to train a model with a new technique: weeks or months - Training is slower than inference when adopting a new technique. PoloQuant implementation time: 31 hours - An engineer at Base 10 implemented the research as a CUDA kernel within 31 hours. Estimated growth in inference engineers demand: 10 to 100 times more in a couple of years - Kiley predicts strong growth in need for inference engineers despite AI coding gains. Classic inference engineer headcount three years ago: a couple hundred - Historical estimate of the field’s size before the generative AI boom. Named entity recognition runtime: 1 millisecond - Example of highly optimized specialized inference runtime at Base 10. Comparable flash/small-model runtime example: 500 milliseconds - Contrast showing the value of specialized optimization for latency-sensitive workloads. Text-to-speech real-time throughput target: about 80 to 100 tokens per second - Used to explain different performance targets by modality. GPU generations discussed: Ampere, Hopper, Blackwell - Kiley describes multiple GPU cycles and their lifecycle in inference. AI-assisted model/workload scale: seven, eight-figure spend behind some workloads - Illustrates the economic scale of optimization opportunities. Open-source runtime concentration: 3 major runtimes - VLLM, SG-Lang, and TensorRT-LLM are identified as the main inference runtimes. Digital copies distributed: 20,000 - Kiley says the book has already had 20,000 digital copies downloaded.

Pivotal Quotes: "inference is, I believe, the most important workload in AI, and it's definitely the stickiest." — Philip Kiley: He explains why Base 10 focused on inference as its wedge into AI infrastructure. "you can't vibe code uptime." — Philip Kiley: He argues that mission-critical inference systems still require human ownership and accountability. "the timeline is often hours." — Philip Kiley: He describes how fast new inference research can move from paper to production.

Implications: Inference is becoming a strategic advantage, not just an infrastructure concern. Teams that understand and control it can lower cost, cut latency, and build better products, especially as agents and multimodal workloads multiply.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast