No Priors
No Priors

State Space Models and Real-time Intelligence with Karan Goel and Albert Gu from Cartesia

This week on No Priors, Sarah Guo and Elad Gil sit down with Karan Goel and Albert Gu from Cartesia. Karan and Albert first met as Stanford AI Lab PhDs, where their lab invented Space Models or SSMs, a fundamental new primitive for training large-scale foundation models. In 2023, they Founded Cartes

Featured Speakers

Albert Gu GuestKaren Goel Guest

Topics Discussed

Episode Summary

Executive Summary: Cartesia co-founders Karen Goel and Albert Gu explain how state-space models (SSMs) like S4 and Mamba challenge transformers on both efficiency and fit for sequential/perceptual data. They describe Sonic, a low-latency text-to-speech product, and argue that the future is hybrid models plus on-device, real-time multimodal AI.

Main Topics: Cartesia and Sonic’s low-latency text-to-speech (Priority: 5/5): The company’s current product, Sonic, is positioned as a fast TTS engine for interactive voice use cases such as gaming and voice agents, where latency is critical. Why SSMs instead of transformers (Priority: 5/5): The founders explain state-space models as compressed, continuously updated memory systems that are naturally suited to sequence modeling and can be more efficient than attention-heavy transformers. Data-type-specific inductive bias (Priority: 4/5): They contrast text, audio, video, and raw signals, arguing that transformers and SSMs each have strengths depending on how compressed or continuous the input is. Hybrid architectures as the practical sweet spot (Priority: 5/5): Both speakers suggest combining mostly SSM layers with a smaller amount of attention for exact retrieval, describing this as a synergistic compromise with strong empirical support. Edge/on-device inference and cheaper intelligence (Priority: 5/5): A major long-term goal is to move inference from the cloud to laptops and edge hardware, enabling real-time AI that is cheaper, faster, and closer to the data source. Speech as an entry point to multimodal models (Priority: 4/5): They frame speech as both a valuable standalone product and a gateway to broader multimodal systems that can ingest audio natively, reason over context, and generate responses intelligently. Research philosophy and aesthetic simplicity (Priority: 3/5): The founders emphasize elegant, simple, fundamental model design as a guiding principle for research and product development, echoing a 'proofs from the book' mindset.

Key Arguments: SSMs are a fundamental alternative to transformers for sequence modeling because they update a compact internal state as new information arrives. For perceptual signals like audio and video, compression is beneficial because the data is highly redundant and streaming-friendly. Transformers are not universally optimal; they rely on tokenization and other assumptions that fit text well but can be unnatural for raw signals. Efficiency matters as much as quality: linear scaling enables constant-time updates per token, unlike transformer attention’s history-dependent cost. A hybrid model with mostly SSMs plus a small amount of attention can outperform either architecture alone by pairing fast state compression with exact retrieval. On-device inference will unlock new applications by reducing latency, compute cost, and orchestration overhead. Speech generation is still far from solved because high-quality TTS requires controllable emotion, role-awareness, and language understanding. Multimodal systems are necessary even to do single modalities well; perfect TTS may require language-level and cross-modal reasoning. Cartesia aims to build infrastructure that makes advanced models cheap enough and small enough to run close to sensors and user devices.

Data Points: Latency reduction in Sonic: ~150 milliseconds - They say Sonic already shaves off about 150 ms versus typical voice systems. Target additional latency savings: ~600 milliseconds - They describe a roadmap to remove another 600 ms over the course of the year. Hybrid layer ratio: ~10:1 - They mention multiple groups finding mostly SSM layers with a small amount of attention to be optimal, around a 10-to-1 ratio. Company headcount: 15 people - They say Cartesia has 15 employees at the time of the conversation. Intern count: 8 interns - They mention a large intern class of eight interns. On-device model size example: ~3 billion parameters - They reference Apple’s on-device models as being around 3B parameters.

Pivotal Quotes: "The dream is to have a game where you have millions of players and they're able to just interact with these models." — Kate: Describing gaming as a major use case for low-latency voice generation. "We think of these state-space models as kind of being fuzzy compressors." — Albert Gu: Explaining the core intuition behind SSMs as compact, continuously updated memory systems. "Would I want to talk to this thing for more than 30 seconds? And if the answer is no, then it's not solved." — Karen Goel: Defining the bar for whether text-to-speech feels truly natural and engaging.

Implications: The episode suggests the next AI wave is about efficient, hybrid, on-device models that reduce latency and cost. For voice, gaming, and edge devices, this could enable more natural, always-available intelligence embedded directly in products.

🔓 Sign Up for Unlimited Episode Search

About No Priors

View all episodes from No Priors