The TWIML AI Podcast
The TWIML AI Podcast

Spiking Neural Nets and ML as a Systems Challenge with Jeff Gehlhaar - TWIML Talk #280

Today we’re joined by Jeff Gehlhaar, VP of Technology and Head of AI Software Platforms at Qualcomm. Qualcomm has a hand in tons of machine learning research and hardware, and in our conversation with Jeff we discuss: • How the various training frameworks fit into the developer experience when worki

Featured Speakers

Jeff Gelhar Guest

Topics Discussed

Episode Summary

Executive Summary: Jeff Gelhar traces Qualcomm’s evolution from wireless systems to AI software platforms, explaining how early work on spiking neural networks shaped today’s edge AI strategy. He argues that the future is a systems problem spanning models, compilers, hardware, and ecosystems, with major opportunities in on-device inference, personalization, TinyML, and cloud-edge collaboration.

Main Topics: Jeff Gelhar’s Qualcomm background and AI arc (Priority: 5/5): Gelhar describes a 30-year Qualcomm career across wireless, systems, software, research, and commercial AI, including a detour into an early machine-learning startup and later leadership of Qualcomm’s AI software commercialization efforts. Spiking neural networks and early biologically inspired AI (Priority: 5/5): He explains Qualcomm’s early research with Brain Corporation and DARPA Synapse-inspired spiking neural networks, why they were attractive for low-power computation, and why training difficulties made them less practical than backprop-based deep learning. Qualcomm’s AI software stack and edge deployment (Priority: 5/5): Gelhar details Qualcomm’s SDKs, GPU/DSP/HTA acceleration, Android NN API support, and TensorFlow Lite backend work, emphasizing low-friction deployment of AI to Snapdragon devices. Framework choice, conversion, and ONNX interoperability (Priority: 4/5): He discusses the proliferation of training frameworks and Qualcomm’s strategy of supporting common ones via converters and ONNX, making deployment easier across TensorFlow, PyTorch, Caffe, and others. Inference, personalization, and privacy-preserving on-device learning (Priority: 4/5): The conversation shifts from generic inference to personalized device experiences, including face/fingerprint unlock, speaker ID, federated learning, and privacy-preserving aggregation for use cases like keyboards and healthcare. Cloud AI 100, edge-to-cloud system design, and 5G (Priority: 4/5): Gelhar says Qualcomm’s AI software approach extends to data center inference and to hybrid systems where edge devices, edge compute, and cloud inference cooperate over low-latency 5G/Wi-Fi links. TinyML and compiler-driven optimization for constrained devices (Priority: 4/5): He highlights TinyML as a push toward microcontroller-scale ML, driven by compact frameworks, quantization, compression, and compiler technology that tailor runtime libraries to specific workloads and hardware.

Key Arguments: Qualcomm’s AI strategy is rooted in long-term systems thinking, not just chips or models, and depends on aligning algorithms, hardware, middleware, and ecosystem APIs. Spiking neural networks were an important exploratory step, but they lacked a principled training method comparable to backprop, limiting practicality for large-scale tasks. Most real-world edge AI now works best by using standard training frameworks and converting models into optimized deployment formats rather than requiring bespoke model development for each chip. ONNX reduces fragmentation by providing an interchange format, making it easier to move models across frameworks and into Qualcomm’s acceleration stack. The next major wave is not just inference, but personalization: devices will adapt to individual users and contexts without relying on massive cloud data. Federated and privacy-preserving learning can improve personalization and timeliness by aggregating signals across users while preserving data privacy. TinyML and compiler specialization will let AI run on extremely small, battery-constrained devices by generating lean runtimes and tailoring operator sets to specific models. Data center inference and edge inference are analogous problems; both need frictionless deployment, high performance, and low power, but they differ in density, operator richness, and system design constraints.

Data Points: Qualcomm tenure: 30 years - Gelhar says he has spent about 30 years at Qualcomm, with a break between tours of duty. Brain Corporation: Small company in San Diego - He identifies Brain Corporation as the company Qualcomm invested in and collaborated with on spiking neural networks. Neurons and power in the brain: 20 billion neurons and 20 watts - Used as an intuition for why low-power biologically inspired computing was appealing. Timeframe of deep learning breakthrough: 2011–2012 - He marks the point when deep neural networks made the case for backprop-based training. On-device learning memory target: 30, 40, 50K of memory - He mentions TinyML ambitions for ultra-small devices. Battery life target: 1–3 years - He describes the goal for very low-power microcontroller-class devices. ONNX operator count: 130+ operators - He cites the broad ONNX operator set when discussing conformance and interoperability. Edge conformance subset: 65 operators - He describes an intended smaller ONNX Edge operator subset for compliance and interoperability.

Pivotal Quotes: "The big theme here is we want to be able to talk about what compliance means, what does it mean to support AI at the edge? How do you make it reduce the frequency? For our customers to do that, right?" — Jeff Gelhar: Summarizing Qualcomm’s ONNX Edge and ecosystem interoperability philosophy. "AI, I view it as a real paradigm shift in computing, and it's a big systems problem." — Jeff Gelhar: His closing argument about why AI requires coordination across hardware, software, models, and ecosystems. "we don't have a fundamentally principled way to train these kinds of systems" — Jeff Gelhar: His explanation for why spiking neural networks did not scale as hoped.

Implications: The interview frames AI’s next phase as practical systems engineering: standardized model interchange, compiler optimization, edge-cloud partitioning, TinyML, and privacy-preserving personalization will matter as much as model quality. Qualcomm is positioning itself as infrastructure for that transition.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast