Episode Summary
Executive Summary: Scott Stevenson argues the voice-AI market has moved beyond standalone speech-to-text, LLM, and TTS components toward integrated perception, understanding, and interaction systems. He says better accuracy, latency, and cost now make voice practical, but the next leap is adaptive, closed-loop, end-to-end agents that learn from use. Deepgram is positioning itself as both model builder and infrastructure provider to power this shift.
Main Topics: Voice AI has matured from novelty to practical interface (Priority: 5/5): Stevenson says audio AI is now good, fast, cheap, and accessible enough to support real products, and that developers worldwide are learning how to build voice applications. He frames talking as a major interface alongside typing and tapping. The future is adaptive, closed-loop learning rather than static models (Priority: 5/5): He argues current models are mostly offline-trained and context-limited, but future systems will update in real time, learn from many conversations, and behave more like federated learning systems that improve over time. Perception, understanding, and interaction should replace the speech/LLM/TTS mental model (Priority: 5/5): Stevenson proposes thinking of voice systems as a pipeline of perception, understanding, and interaction, potentially all fused into one system depending on the task, rather than separate speech-to-text, LLM, and TTS products. Deepgram’s strategy: build both models and infrastructure (Priority: 4/5): He explains Deepgram combines foundational model development with hosted infrastructure because voice systems are too complex to leave only to generic infrastructure or generic model providers, especially at scale. Voice agents are already viable for targeted use cases (Priority: 4/5): Stevenson highlights call centers, after-hours business phone coverage, healthcare support, food ordering, and internal productivity tools as near-term wins for voice agents, especially where the task is simple and the workflow is bounded. Real-time voice interaction is close, but not fully human-like yet (Priority: 4/5): He says latency budgets of roughly 250-750 ms are now achievable at reasonable cost, but current systems can still be tripped up by edge cases, mispronunciations, and adversarial or unusual inputs.
Key Arguments: Voice is becoming a primary human-machine modality because accuracy, speed, cost, and accessibility have finally crossed a useful threshold. The biggest shift ahead is from static supervised models to systems that learn continuously from live interactions and improve over time. Audio systems should be designed as perception, understanding, and interaction layers, not just speech-to-text plus LLM plus TTS. Text is a useful debugging and control layer, but it throws away rich acoustic context such as tone, distance, background noise, and emotion. Voice agents must be optimized for specific use cases and cost constraints; a single giant model will be too expensive for many deployments. Deepgram believes model quality improves most when tightly coupled with customer use cases, infrastructure, and feedback loops. Diarization lags other tasks partly because buyers are less willing to pay for it, so model investment follows economics as much as technical possibility. Voice will not replace typing and screens entirely, but it will become a favored and highly productive mode for many tasks.
Data Points: Time since last podcast appearance: about 7 years - Sam notes it has been seven years since Scott last appeared on the show. Company size: about 100 people - Stevenson says Deepgram has grown into a roughly hundred-person company. B2B customer count: over 400 customers - Deepgram serves audio AI models to more than 400 B2B customers. Historical phone-call speech accuracy: about 50% - He says speech systems seven years ago were roughly every other word incorrect on low-quality phone audio. Perception time budget: 250 milliseconds or less - He says perception in a real-time voice system needs to happen within about 250 ms. Total response budget: 500 to 750 milliseconds - He estimates the full perception-understanding-generation loop can fit inside roughly half to three-quarters of a second. Reasonable voice-agent operating cost: $1 to $5 an hour - He says current systems can hit the latency target at this cost range. Traditional call-center hourly wage: $2 to $5 an hour - He compares voice-agent operating cost with call-center labor in the Philippines or India. Conversation context horizon today: 5 to 120 minutes - He notes prompts can grow over time, but many systems lose context after longer conversations, especially around one to two hours. Productivity horizon: next 1 to 2 years - He predicts major gains from voice assistance and note-taking within this timeframe.
Pivotal Quotes: "We shouldn't be talking about speech attacks models anymore, LLM models, or TTS models. We should be talking about perception models, understanding and interaction models." — Scott Stevenson: He reframes the architecture of voice AI around functional layers rather than separate product categories. "The next version will be it can update and learn throughout time." — Scott Stevenson: He describes the shift from offline-trained systems to continuously learning models. "Voice is back. You know, it's back again." — Scott Stevenson: He emphasizes that voice is re-emerging as a major human-computer interface after years of being secondary to text.
Implications: Voice AI is moving from demo to deployment. Expect more practical agents in call centers, support, and productivity workflows first, then deeper end-to-end systems that learn continuously and reduce the need to force everything through text.