The TWIML AI Podcast
The TWIML AI Podcast

Open Source at Qualcomm AI Research with Jeff Gehlhaar and Zahra Koochak - #414

Today we're joined by Jeff Gehlhaar, VP of Technology at Qualcomm, and Zahra Koochak, Staff Machine Learning Engineer at Qualcomm AI Research. If you haven’t had a chance to listen to our first interview with Jeff, I encourage you to check it out here! In this conversation, we catch up with Jef

Featured Speakers

Jeff Gelhar GuestZara Kuchek Guest

Topics Discussed

Episode Summary

Executive Summary: Qualcomm’s Jeff Gelhar and Zara Kuchek describe how the company is pushing AI across devices and cloud through a systems-level approach that combines hardware, software, compilers, and research. They cover Snapdragon AI engine progress, open-sourced AIMET for quantization/compression, TVM and MLIR compiler work, and federated learning research aimed at privacy, efficiency, and real-world edge deployment.

Main Topics: Qualcomm’s systems-level AI strategy (Priority: 5/5): Jeff frames Qualcomm AI as a cross-layer effort: hardware, software, toolkits, and algorithms must be optimized together to deliver efficient on-device AI across mobile, automotive, robotics, IoT, and servers. Snapdragon AI engine and developer ecosystem (Priority: 5/5): The discussion highlights the Snapdragon 865, fifth-generation AI engine, and expanding support through Android NNAPI, Hexagon NN Direct, and Snapdragon NPE SDK to make acceleration easier for developers and OEMs. AIMET open source toolkit for model efficiency (Priority: 5/5): Zara explains Qualcomm’s AI Model Efficiency Toolkit, including quantization and compression methods, and how it is integrated as plugins for TensorFlow and PyTorch to fit existing workflows. Compiler infrastructure: TVM and MLIR (Priority: 4/5): Jeff discusses compiler-based approaches to generating optimized code for Hexagon, contributing a Hexagon backend to TVM, and exploring MLIR as a future path for IR-to-IR translation and deployment efficiency. Federated learning research for edge privacy and efficiency (Priority: 4/5): Zara outlines current research on federated learning: handling heterogeneous users/devices, improving privacy, reducing communication, and lowering memory requirements while keeping training on device. Open source and industry partnerships (Priority: 4/5): Both speakers emphasize collaboration with Google, Microsoft, the TVM community, and other partners, using open source and ecosystem alliances to accelerate adoption and productization. Productization timeline and research-to-market pipeline (Priority: 4/5): Jeff explains that research often lands in commercial products years later, with techniques moving from papers into toolkits, partner releases, and eventually broader open source or hardware generations.

Key Arguments: AI on the edge is a systems problem, not just a model problem; performance and efficiency require coordination across hardware, compilers, software stacks, and algorithms. Qualcomm’s AI engine evolution expands hardware acceleration to more operators, enabling better performance and lower power on mobile and other edge devices. Developers need multiple abstraction levels: high-level APIs for ease of use and lower-level APIs like Hexagon NN Direct for highly optimized workloads. AIMET is designed to work as a plugin inside TensorFlow and PyTorch rather than a separate workflow, making compression and quantization accessible to existing developers. Quantization and compression can be improved through techniques like cross-layer equalization, bias correction, channel pruning, and SVD-based tensor decomposition. Compiler frameworks like TVM can generate machine code tailored to target hardware, enabling whole-graph optimization and better exploitation of cache, vector units, and memory movement. Federated learning is valuable because it keeps data on-device, improving privacy while enabling collaborative model training across distributed devices. Research on federated learning must address heterogeneous data quality, communication cost, and memory limits before it can become broadly commercial. Open source work is used strategically to influence ecosystems, build developer trust, and accelerate adoption of Qualcomm-targeted optimization paths. Qualcomm sees a continuum from tinyML to megaML, spanning sensors, phones, cars, edge servers, and cloud infrastructure under one AI/connectivity umbrella.

Data Points: Snapdragon AI engine generation: 5th generation announced; 6th generation planned later in the year - Jeff describes Qualcomm’s AI engine roadmap and next product cycle Android Neural Networks support: Version 1.2 - Expanded support with more accelerated operators on Qualcomm hardware Google ASR performance improvement: 3x to 5x faster - Google ASR running on Hexagon with Qualcomm partnership Google ASR power reduction: 30% reduction - Same Hexagon deployment example for speech recognition Federated learning memory reduction: 90% less memory requirement - Zara describes trading computation for lower memory use in training Federated learning compute tradeoff: 30% more computation - Same memory-saving technique in federated learning research AI engine announcement timing: December of last year / later this year - Snapdragon 865 and fifth-gen engine announced in December; sixth-gen expected later in year Lab-scale federated learning evaluation: ~1,000 devices - Jeff describes planned real-world prototyping at scale

Pivotal Quotes: "you cannot look at the hardware in isolation or the software in isolation or the bit widths in isolation" — Jeff Gelhar: Explaining Qualcomm’s systems-level approach to efficient AI deployment "we designed this as plug-ins to TensorFlow and PyTorch" — Jeff Gelhar: Describing how AIMET fits into existing developer workflows "we distribute this training between the edge devices" — Zara Kuchek: Defining federated learning and why it improves privacy by keeping data local

Implications: Qualcomm is turning edge AI into an integrated stack problem, which should improve real-device performance, privacy, and developer adoption. Expect more compiler-driven optimization, federated learning pilots, and AI features to surface first in partner products, then broader releases.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast