Super Data Science: ML & AI Podcast with Jon Krohn
Super Data Science: ML & AI Podcast with Jon Krohn

939: Mixture-of-Experts and State-Space Models on Edge Devices, with Tyler Cox and Shirish Gupta

State space models (SSMs), granite models, and Mamba: Dell’s Tyler Cox and Shirish Gupta discuss with Jon Krohn why state space models can process information so efficiently, and how Dell’s AI factory helps enterprises manage custom AI workloads. Hear the latest on the Dell Pro AI Studio and Dell’s

Featured Speakers

Jon Krohn HostSharish Gupta GuestTyler Cox Guest

Topics Discussed

Episode Summary

Executive Summary: Dell’s Tyler Cox and Sharish Gupta explained how Dell’s AI PC solution (formerly Dell Pro AI Studio, now being rebranded under Dell AI Factory) makes local on-device AI practical for enterprises by abstracting hardware complexity, simplifying deployment, and adding fleet-wide manageability. They also detailed Granite 4.0’s hybrid transformer + state-space and mixture-of-experts design, showing how these architectures improve long-context efficiency and make advanced AI viable on PCs, edge devices, and workstations.

Main Topics: Dell AI Factory / AI PC solution rebrand and scope (Priority: 5/5): The guests described the evolution of Dell Pro AI Studio into a broader AI Factory offering that spans the PC, edge, and server ecosystem, with modular components for models, frameworks, and management. Why local and edge AI matters (Priority: 5/5): They argued that local inference reduces latency, protects privacy, supports offline/always-on use cases, and helps control cost as AI usage scales across enterprises. Technical architecture for enterprise deployment (Priority: 5/5): Tyler outlined how Dell abstracts away heterogeneous silicon, runtimes, quantization, and model-management complexity while preserving familiar APIs and enterprise-grade lifecycle controls. Granite 4.0: hybrid state-space + transformer models (Priority: 5/5): The conversation covered IBM’s Granite 4 family, especially its hybrid design using Mamba/state-space layers with transformers to improve long-context performance and memory efficiency. Mixture of experts for efficient scaling (Priority: 4/5): They explained how sparse MoE models activate only subsets of parameters per token/request, enabling large-model capability with smaller-model-like throughput costs. Enterprise AI failure modes and best practices (Priority: 4/5): The guests discussed why many enterprise AI projects fail and emphasized clear objectives, human-in-the-loop oversight, using the right tool for the job, and strong manageability. Hardware benchmarks and AI PC gains (Priority: 4/5): Sharish highlighted performance, battery, and AI inference gains from newer AI PC chipsets versus older non-AI systems to show practical benefits of adopting modern hardware.

Key Arguments: Local inference is valuable because not all workloads need frontier-scale cloud models; many tasks can run faster, cheaper, and more privately on-device. Enterprises need an AI deployment stack that is as easy as cloud APIs but with the governance, lifecycle management, and security expected of enterprise software. Dell’s approach is to hide low-level complexity such as silicon diversity, execution providers, and quantization details behind a consistent API and management layer. A hybrid model architecture can outperform pure transformers on long-context tasks by improving memory efficiency and reducing quadratic scaling pressure. Mixture-of-experts models are especially useful because they preserve the quality of large models while using only a subset of parameters at inference time. Hybrid, on-device, and cloud inference will coexist; the right compute target should depend on the task, latency needs, privacy constraints, and cost. Enterprise AI projects fail when teams use GenAI indiscriminately, lack clear success criteria, or ignore production lifecycle management and human oversight. AI PCs and modern chipsets deliver meaningful performance and battery improvements, making them a viable part of enterprise AI strategy rather than a niche option.

Data Points: Podcast episode reference: Episode 921 - Sharish Gupta’s previous appearance was cited as one of the podcast’s most watched YouTube episodes YouTube views: Over 100,000 views - Episode 921 was noted as surpassing 100k views on YouTube Inference workload forecast: 90% - Sharish said almost 90% of AI workloads are expected to be inference by the turn of the decade Cloud latency avoided: ~200 ms - Sharish referenced roughly 200 ms of cloud latency that local inference can remove Model catalog scale: 2 million models - Tyler noted Hugging Face has around 2 million models available in the ecosystem Granite 4 micro context scaling example: 8 sessions at 128K context - A comparison was given for memory usage in long-context settings on the Granite 4-H micro model Memory usage comparison: 15 GB vs 80 GB - Eight 128K sessions on Granite 4-H micro used about 15 GB versus about 80 GB for a pure transformer architecture Granite 4 family sizes: 3B, 7B, 32B parameters - Tyler described Granite 4 Micro as 3B, Granite 4-H Tiny as 7B, and Granite 4-H Small as 32B total parameters Active parameters in MoE models: 1B active / 9B active - Granite 4-H Tiny has about 1B active parameters; Granite 4-H Small has about 9B active parameters IBM benchmark: Top five on BFCL v3 - Tyler said Granite 4-H Small ranks in the top five on Berkeley Function Calling Leaderboard v3 Lens of transformer scaling: Quadratic memory growth - Used to contrast transformers with state-space models' linear scaling behavior Dell AI PC benchmark: 88% more battery runtime - Sharish cited Dell Pro Plus with Intel Core Ultra 200V/Lunar Lake for Microsoft Teams meetings Graphics performance: 4.8x higher - Dell Pro Plus with Intel Lunar Lake vs. non-AI PC generation On-device AI CPU performance: ~10x higher - Sharish described CPU gains on the newer AI PC versus a non-AI PC On-device AI GPU performance: 4.2x higher - Sharish described GPU gains on the newer AI PC versus a non-AI PC Text generation performance: >6x higher - Sharish cited greater than six times higher performance on Mistral 7B and Llama 3.1 for text generation Image generation performance: Up to 5.6x higher - Sharish cited Stable Diffusion 1.5 performance gains on the newer AI PC

Pivotal Quotes: "We believe, I firmly believe that the future of AI is hybrid, right? And workloads will run seamlessly across the cloud, the edge, and the PC fleet." — Sharish Gupta: Explaining why local, edge, and cloud inference will coexist as part of enterprise AI strategy "So what we're doing is they're picking portions of the model dynamically for each request, and really at a token level to answer and kind of retain only the most important pieces of the model in answering a given request." — Tyler Cox: Describing how mixture-of-experts models achieve efficiency while keeping model capability high "It's not enough to just bring a model and run it locally on one PC. You need to have the ability to deploy and manage that through the entire lifecycle of that AI workload or app across a whole fleet of PCs." — Sharish Gupta: Explaining the enterprise requirement for lifecycle management, not just local execution

Implications: The episode frames AI PCs as a serious enterprise platform, not a novelty: they can cut latency, protect data, and reduce costs while preserving manageability. Expect more hybrid architectures, more edge inference, and faster enterprise AI rollouts.

🔓 Sign Up for Unlimited Episode Search

About Super Data Science: ML & AI Podcast with Jon Krohn

View all episodes from Super Data Science: ML & AI Podcast with Jon Krohn