No Priors
No Priors

The Future of Voice AI: Agents, Dubbing, and Real-Time Translation with ElevenLabs Co-Founder Mati Staniszewski

Imagine learning chess from a grand master, or negotiating tactics from an expert FBI hostage negotiator. ElevenLabs’ voice AI technology is making that unlock possible. Sarah Guo sits down with Mati Staniszewski, co-founder of ElevenLabs, to explore how the three-year old company is transforming ho

Featured Speakers

Maddie Staniszewski Guest

Topics Discussed

Episode Summary

Executive Summary: Maddie Staniszewski explains how ElevenLabs evolved from a voice generation startup into a broader audio AI platform spanning creative tools and enterprise agents. The conversation covers why voice is becoming a core human-computer interface, how the company sequences research and product, where voice quality and benchmarks still fall short, and why education, customer experience, dubbing, and immersive media may be the biggest future applications.

Main Topics: ElevenLabs’ mission and product scope (Priority: 5/5): ElevenLabs builds foundational audio models and layered products for narration, dubbing, customer experience, personal AI, and immersive media, all aimed at improving how humans interact with technology through voice. From research lab to platform company (Priority: 5/5): Maddie describes the company’s structure as a set of focused labs: first voice, then agent orchestration, then adjacent areas like music, with research and product developed in parallel around specific problems. Why voice matters as the next interface (Priority: 5/5): The company’s founding insight is that speech is the most natural interface and will increasingly replace keyboard/screen interaction across phones, computers, robots, and global communication. Quality, evaluation, and voice selection (Priority: 4/5): A major challenge is that audio quality is highly subjective and hard to benchmark. ElevenLabs addresses this with voice experts, custom guidance, and better audio labeling/evaluation workflows. Enterprise use cases and business model (Priority: 4/5): On the enterprise side, the company focuses on customer support, proactive commerce, internal training, and multilingual experiences, combining platform software with forward-deployed engineering. Competitive differentiation versus foundation model labs (Priority: 4/5): Maddie argues that large labs may not prioritize the product layer and specialized audio focus, while ElevenLabs’ advantage comes from audio-specific research, controllability, and shipping production-ready workflows. Future applications: education, companions, and real-time translation (Priority: 5/5): The most exciting long-term opportunities are personalized AI tutors, more natural assistant/companion experiences, and real-time multilingual communication that preserves voice and emotion.

Key Arguments: ElevenLabs started with a specific pain point—poor-quality dubbing and narration—and expanded from there into a broader audio platform. The company believes voice is the most natural interface for interacting with technology and will become increasingly important across devices and contexts. Research and product should not be sequentially isolated; ElevenLabs creates research-specific labs and ships product in parallel so customer value arrives faster. Audio is unusually hard to evaluate because perception depends not just on speech quality but on voice fit, emotion, accent, and use case. Enterprises need both model capability and implementation help; ElevenLabs differentiates by combining platform software with forward-deployed engineering and voice expertise. The long-term moat is less the base model alone and more the ecosystem: workflows, integrations, distribution, and collected voices. General-purpose model companies may not optimize deeply for the product layer or the specialized controllability that voice applications require. Education is likely the biggest underdeveloped voice use case, with AI tutors and interactive learning experiences poised to scale. AI companions will exist, but Maddie is more excited about a Jarvis-style super-assistant than social companion use cases. Voice and agentic systems will eventually support real-time translation, dubbing, and multimodal interaction across audio, text, image, and video.

Data Points: Company size: 350 people - ElevenLabs’ global headcount Annual recurring revenue: $300 million ARR - Current company scale Business mix: ~50/50 self-serve and enterprise - Revenue split between creative platform subscriptions and enterprise agents Monthly active users: 5 million monthly actives - Creative platform usage Enterprise customers: A few thousand customers - Includes Fortune 500s and fast-growing AI startups Languages covered by STT: Almost 100 languages - Speech-to-text coverage mentioned for creative and transcription products Scribe V2 latency: Under 150 milliseconds - Recent speech-to-text model performance Scribe V2 accuracy: 93.5% accuracy - Reported on Flares across top 30 languages Model benchmark comparison: Beating top models on benchmarks - Claim made for speech-to-text / orchestration performance Research team estimate: 10 researchers - Maddie says ElevenLabs has roughly 10 of the 50-100 top audio researchers Research advantage window: 6–12 months - Head start from research that translates into customer value Build-vs-wait threshold: About 3 months - If a product change will take longer than this, they build it

Pivotal Quotes: "voice is the interface of the future" — Maddie Staniszewski: Explaining the company’s core thesis about human-computer interaction "research is head start" — Maddie Staniszewski: Describing how ElevenLabs thinks about the relationship between models, product, and long-term advantage "I think education with voice... is just going to be such a big thing" — Maddie Staniszewski: Answering what future use case will most change content consumption and learning

Implications: Voice AI is moving from novelty to infrastructure. Winners will pair strong models with productization, integrations, and domain-specific workflows, especially in customer support, education, and real-time multilingual communication.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors