The TWIML AI Podcast
The TWIML AI Podcast

Stable Diffusion and LLMs at the Edge with Jilei Hou - #633

Today we’re joined by Jilei Hou, a VP of Engineering at Qualcomm Technologies. In our conversation with Jilei, we focus on the emergence of generative AI, and how they've worked towards providing these models for use on edge devices. We explore how the distribution of models on devices can help

Featured Speakers

Gile Ho Guest

Topics Discussed

Episode Summary

Executive Summary: Sam Charrington interviews Gile Ho, VP of Engineering at Qualcomm Technologies, about Qualcomm AI Research and the push to run generative AI on-device. Ho traces the group’s roots in information theory, explains why compression and quantization are central to efficient edge AI, and details Qualcomm’s progress on stable diffusion, LLMs, multimodal systems, and hybrid edge-cloud inference.

Main Topics: Information theory as the foundation for modern AI (Priority: 4/5): Ho describes his background in information theory and signal processing and argues that foundational concepts like KL divergence show the deep connection between Shannon-style information theory and contemporary machine learning. Qualcomm AI Research mission and focus (Priority: 5/5): He explains that Qualcomm AI Research was formed around 2018 after AI startup acquisitions, with an agenda centered on power efficiency for on-device AI and personalization, especially for speech, audio, vision, and wireless use cases. Quantization and compression as enabling technologies (Priority: 5/5): Ho emphasizes quantization research as a major differentiator for edge deployment, highlighting its role in making models smaller, faster, and more power-efficient, and linking it to Qualcomm’s broader history in data compression. Stable diffusion on device (Priority: 5/5): The discussion covers why running stable diffusion locally matters—privacy, cost, and reliability—and how Qualcomm tackled the challenge through full-stack optimization, especially post-training quantization and the AddRound method. Differences between vision generative models and LLMs (Priority: 4/5): Ho contrasts LVMs/diffusion models and LLMs, arguing that both are transformer-heavy but differ in compute characteristics, output representations, and quantization opportunities; LLMs can often be compressed more aggressively than pixel-based models. Multimodal and System 2 AI (Priority: 4/5): He describes Qualcomm’s System 2 research track and the push toward combining vision and language for more cognitive, reasoning-oriented interactions, including visual prompting and benchmark-driven progress in visual understanding. Future priorities: multi-token generation, training efficiency, and hybrid AI (Priority: 5/5): Ho says Qualcomm is reorganizing around generative AI while still supporting non-generative work, and identifies three priorities: multi-token generation, more efficient training of large models, and hybrid edge-cloud AI architectures.

Key Arguments: Information theory and machine learning are deeply connected, with concepts like KL divergence originating in information theory and later becoming central to neural network training. Qualcomm AI Research was intentionally built around on-device efficiency and personalization, not just model performance, because mobile and edge constraints demand power-aware AI. Quantization is not just an implementation detail; it is a core research lever that enables large models to run on-device with acceptable latency and energy use. Stable diffusion on-device is valuable because it preserves privacy, lowers cloud inference cost, and improves reliability when cloud capacity is unavailable. The main technical obstacles for stable diffusion were model size and inference latency, which required end-to-end optimization across models, software, and silicon. Advanced post-training quantization methods like AddRound are necessary because naive quantization harms quality too much; localized optimization improves signal-to-noise and practical output fidelity. The same techniques that work for diffusion models are broadly transferable to other vision-language models, though LLMs require different handling because their outputs are token-based rather than pixel-based. LLMs are inherently inefficient in their current inference form because very large models generate one token at a time, so multi-token generation could substantially reduce bandwidth and cost. Future AI systems will likely be hybrid, with edge devices handling substantial local inference and the cloud reserved for harder or heavier tasks.

Data Points: Years since PhD: ~20 years - Ho says he earned his PhD from UCSD about 20 years ago before joining Qualcomm. Qualcomm AI Research launch timing: Around 2018 - He says the AI research initiative began after Qualcomm acquired several AI startups. Team size: 200 to 300 people - Ho estimates the Qualcomm AI Research group is now in this range. Internet traffic from images/video: 80% to 90% - He cites image and video as the dominant share of internet traffic to motivate compression research. Stable diffusion model size: 1.1 billion parameters - Used as a benchmark for the challenge of running generative models on-device. Typical on-device model size before stable diffusion: Less than 100 million parameters - Ho contrasts prior on-device workloads with stable diffusion’s much larger scale. Denoising steps: 20 to 50 steps - He notes stable diffusion requires iterative denoising, adding to runtime complexity. On-device demo latency: Less than 15 seconds - Qualcomm achieved a stable diffusion demo running on device in this time frame. Best reported latency: Around 13 seconds - Ho gives the current record for on-device stable diffusion performance. Target latency improvement: Less than 5 seconds - The team hopes to reach this within the next few months. Quantization format for diffusion demo: 8-bit weights and 16-bit activations - He cites the data format used for the MWC demo of stable diffusion. Optimization time for AddRound: Hours - Current AddRound optimization for stable diffusion takes hours. Optimization time goal: Minutes - The team aims to reduce AddRound optimization from hours to minutes. ChatGPT model size: 175 billion parameters - Used as an example of why current LLM inference is inefficient. LLM edge model example: 7 billion parameters - He mentions popular edge LLMs such as Llama starting around this scale.

Pivotal Quotes: "there's a kind of a match made in heaven, so to speak, between information theory and artificial intelligence" — Gile Ho: Ho explains the conceptual link between his academic background and modern ML. "privacy coming down to fix" — Gile Ho: He argues that running stable diffusion on-device keeps prompts and content local, protecting privacy. "running the general AI models actually is very inefficient" — Gile Ho: He motivates Qualcomm’s push toward multi-token generation and hybrid AI architectures.

Implications: Edge AI is moving from simple inference to large generative workloads. For consumers this means more private, faster experiences; for industry it implies new compression, quantization, and hybrid-cloud architectures will be key competitive advantages.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast