The TWIML AI Podcast
The TWIML AI Podcast

Simplifying On-Device AI for Developers with Siddhika Nevrekar - #697

Today, we're joined by Siddhika Nevrekar, AI Hub head at Qualcomm Technologies, to discuss on-device AI and how to make it easier for developers to take advantage of device capabilities. We unpack the motivations for AI engineers to move model inference from the cloud to local devices, and expl

Featured Speakers

Siddhika Nevrekar Guest

Topics Discussed

Episode Summary

Executive Summary: The conversation explores the rise of on-device AI and Qualcomm’s AI Hub, focusing on why developers still struggle to deploy models locally despite growing demand. Siddhika Nevrekar explains that cost, privacy, connectivity, latency, and fragmented hardware/software stacks are driving on-device adoption, while AI Hub aims to simplify model conversion, benchmarking, and device testing across many devices in minutes.

Main Topics: Career path from cloud ML to on-device AI (Priority: 4/5): Siddhika traces her journey from Microsoft’s cloud-scale search and ranking systems to Apple’s device ML work, where she realized on-device AI was real and transformative, then to founding Tetra AI and joining Qualcomm. Why on-device AI matters now (Priority: 5/5): She argues that developers move models on-device to reduce cloud costs, improve privacy, ensure offline connectivity, and achieve lower latency by using the user’s own device as the compute platform. Hardware stack fragmentation and developer pain (Priority: 5/5): The discussion details the complexity of targeting CPUs, GPUs, NPUs/neural engines, runtimes, OS versions, and device generations, which creates a steep learning curve and large testing matrix for developers. Performance limits: compute vs memory (Priority: 5/5): Even though modern SoCs and neural processors can execute AI workloads quickly, memory constraints—especially for large language models—become the bottleneck on mobile devices, requiring compression, quantization, and model partitioning. AI Hub as a developer abstraction layer (Priority: 5/5): Qualcomm’s AI Hub is presented as a service where developers upload a PyTorch model, choose target devices, and receive answers about whether the model runs, how well it performs, and what runtime/compilation path to use. Testing, benchmarking, and power measurement (Priority: 4/5): Testing is described as one of the biggest barriers to adoption; AI Hub helps with compatibility and speed metrics, while power measurement remains more complex due to environmental variables and device-state dependencies. Future directions: convergence, but continued innovation (Priority: 4/5): Vision models are relatively converged today, but ML continues to advance into multimodal systems (vision, audio, language, editing), meaning the ecosystem will keep evolving and new device challenges will appear.

Key Arguments: On-device AI is driven by practical business and user needs: lower cloud spend, better privacy, offline availability, and faster experiences. Developers face major friction because each device, OS, runtime, and chip generation introduces a different compatibility and optimization problem. Modern AI chips are highly capable, but memory is increasingly the limiting factor when running large models locally, especially on mobile. The software ecosystem is fragmented across PyTorch, TF Lite, ONNX, DirectML, and vendor-specific runtimes, so there is no universal deployment path. Successful deployment depends on close coordination between chip manufacturers and runtime ecosystems so that operators and kernels are supported natively. AI Hub is intended to reduce the “where do I even start?” problem by turning model deployment into a fast validation workflow with device selection and performance feedback. Testing on real devices remains necessary because virtual environments cannot fully capture performance, compatibility, and power behavior. Power measurement is harder than latency or accuracy because it depends on thermal conditions, background app load, and lab-grade setups. Vision workloads are relatively mature and easier to deploy today, but multimodal and novel operators will continue to stress the stack. IoT and automotive have different form-factor constraints, but the same software principles apply; the biggest differences are power, connectivity, and control over the full stack.

Data Points: Developer interviews: 100 developers - Qualcomm spoke to 100 developers to understand pain points in on-device AI adoption. AI Hub model-to-answer promise: 5 lines of code and 5 minutes - AI Hub is designed to tell a developer whether a PyTorch model works on a device and how well it runs very quickly. Time-to-first-token / inference: 10s and 20s tokens - On Snapdragon X Elite laptop experiences, once the first token is generated, models can continue at roughly 10s to 20s tokens (as stated in the transcript). Apple tenure: 6 years - Siddhika spent six years at Apple working on device ML use cases such as Face ID and watch detection. Microsoft tenure: 9 years - She spent about nine years at Microsoft working on cloud ML, Bing, Bing Maps, Bing Shopping, and ranking/relevance.

Pivotal Quotes: "I did not believe ML on device existed." — Siddhika Nevrekar: She describes her initial skepticism before working on Apple Neural Engine and device ML use cases. "What if, as a developer, I could come with a PyTosh model... Can you give me an answer whether this model runs?" — Siddhika Nevrekar: She explains the foundational idea behind AI Hub: simplified model validation across target devices. "This is where ML is going to always surprise us. It doesn't stop innovating." — Host (intro framing): The episode opens by highlighting the rapid expansion from vision to multimodal and generative on-device experiences.

Implications: On-device AI is becoming a mainstream necessity, but adoption depends on simplifying deployment across fragmented hardware and runtimes. Tools like AI Hub could unlock broader developer participation and accelerate private, offline, low-latency AI experiences.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast