The TWIML AI Podcast
The TWIML AI Podcast

From Particle Physics to Audio AI with Scott Stephenson - TWiML Talk #19

This week my guest is Scott Stephenson. Scott is co-Founder & CEO of Deepgram, which has developed an AI-based platform for indexing and searching audio and video. Scott and I cover a ton of interesting topics including applying machine learning techniques to particle physics, his time in a lab

Featured Speakers

Scott Stevenson Guest

Topics Discussed

Episode Summary

Executive Summary: Scott Stevenson explains how Deepgram grew out of particle physics work on dark matter detection, where signal extraction from noisy waveforms translated naturally to audio AI. He describes Deepgram’s non-text-based audio search approach, its use cases in fraud and call-center QA, and the open-source Kerr framework designed to make deep learning model development more declarative and efficient.

Main Topics: From particle physics to audio AI (Priority: 5/5): Stevenson traces Deepgram’s founding team back to graduate work in dark matter experiments, where they learned to extract signals from waveform data using machine learning techniques that closely resemble audio processing. Deepgram’s audio search without text as the intermediary (Priority: 5/5): Deepgram does not rely on speech-to-text as the primary representation; instead it builds indexes from deep neural network activations to enable direct search over audio content and topics. Audio use cases and market fit (Priority: 5/5): The conversation highlights practical applications such as searching podcast/audio archives, customer service analytics, fraud detection, compliance, and QA for large call-center datasets. Why physics backgrounds show up in AI (Priority: 4/5): Stevenson argues that physics and AI are both fundamentally about extracting structure from noisy data, which is why many physicists transition naturally into machine learning work. China lab story and rapid experimental execution (Priority: 3/5): He describes how the PandaX dark matter experiment was enabled by a massive underground lab built in western China in under nine months, illustrating a highly execution-oriented environment. Kerr: an open-source declarative deep learning framework (Priority: 5/5): Deepgram released Kerr to simplify and accelerate model experimentation by specifying architectures in YAML/JSON, while abstracting over TensorFlow, Theano, and PyTorch backends.

Key Arguments: Particle-physics workflows and audio AI both involve extracting meaningful signals from noisy waveform data, making the transition between domains unusually natural. Deepgram’s core innovation is indexing audio directly from neural network representations rather than converting everything to text first. Traditional speech-to-text performs poorly on low-quality, real-world audio, which is exactly where many valuable datasets exist. Audio search is especially useful for large-scale business call recordings, where humans cannot manually review millions of interactions. Fraud detection and quality assurance/compliance are among the strongest enterprise use cases because they rely on pattern detection across massive call volumes. Physics-trained researchers are overrepresented in ML partly because both fields are about finding trends, outliers, and causal structure in data. Kerr was built because existing deep-learning frameworks were too slow and cumbersome for rapid experimentation at Deepgram’s scale and team size. Open sourcing Kerr was both a community contribution and a recruiting strategy, leveraging external contributions and increasing visibility. The company believes the broader world is moving toward an expectation that all audio and media content should be searchable. For niche classification problems, Deepgram can use transfer learning with customer-labeled audio to train topic models on top of its general audio-search capabilities.

Data Points: Depth of underground lab: 2 miles underground - The PandaX dark matter experiment was housed in what Stevenson described as the deepest lab in the world in western China. Experiment team size: Around two dozen people - He described the dark matter experiment as built by roughly two dozen collaborators. China lab build time: Less than 9 months - The underground lab was reportedly blasted out and converted into a working research facility very quickly after approval. PandaX experiment duration: Under 4 years - Stevenson said the experiment was designed, built, deployed, and analyzed within less than four years. Deepgram improvement in search accuracy: From about 20% to 80-90% - He said early Deepgram prototypes improved audio search from very poor accuracy to strong retrieval performance within 4-5 months. Deepgram team size: 8 people - He referenced Deepgram as a small engineering team when discussing architectural tradeoffs. Labeled files needed for useful topic modeling: 50-100 labeled files - He estimated this as enough to begin getting useful consumer-facing topic-search results. Labeled files for ~90% accuracy: Around 1,000 labeled files - He said larger labeled datasets can push topic modeling toward about 90% accuracy. Call review practice today: 1% to 5% of calls - He described current enterprise QA/compliance workflows where humans inspect only a small fraction of calls. Audio file granularity: 10-minute to hour-long files - He noted that training examples are generally in this range, depending on the task.

Pivotal Quotes: "AI is kind of just information physics." — Scott Stevenson: He used this to explain why physicists often transition well into machine learning and audio AI. "We don’t want to think about Python code and how do we put all of that in there." — Scott Stevenson: He was describing why Deepgram built Kerr as a declarative model-definition layer. "We’re not building an index out of text anymore. We’re building an index, you know, like actually out of activations in a deep neural network." — Scott Stevenson: This captured Deepgram’s key technical distinction from standard speech-to-text search pipelines.

Implications: The episode suggests audio AI is moving beyond transcription toward direct semantic search and operational intelligence. For enterprises, this could transform QA, compliance, and fraud detection; for consumers, it points to a future where spoken media is as searchable as text.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast