The TWIML AI Podcast
The TWIML AI Podcast

Generative AI at the Edge with Vinesh Sukumar - #623

Today we’re joined by Vinesh Sukumar, a senior director and head of AI/ML product management at Qualcomm Technologies. In our conversation with Vinesh, we explore how mobile and automotive devices have different requirements for AI models and how their AI stack helps developers create complex models

Featured Speakers

Vinesh Sukumar Guest

Topics Discussed

Episode Summary

Executive Summary: Vinesh Sukumar of Qualcomm describes how AI at the edge is shifting from CNN-heavy image use cases toward transformers, multimodal, and generative workloads. He explains Qualcomm’s strategy across hardware and software: fixed-point compute, efficient memory/bandwidth design, quantization, and a unified AI stack spanning mobile, automotive, PC, and enterprise. The conversation emphasizes edge-specific MLOps, personalization, hybrid cloud/edge execution, and micro-tile inferencing for faster, lower-power model execution.

Main Topics: Evolution of edge AI use cases (Priority: 5/5): Edge AI is expanding beyond image/video tasks into text, multimodal, commerce, recommendation, and generative AI, driving a shift from CNN-centric designs toward transformer-friendly architectures. Hardware/software co-design for edge performance (Priority: 5/5): Qualcomm’s approach balances compute, memory, bandwidth, data types, and compiler/runtime optimizations to meet KPIs like latency, power, quality of service, and footprint. Mobile vs. automotive AI requirements (Priority: 5/5): Mobile emphasizes low power and small memory for camera and text-based experiences, while automotive requires high concurrency, larger tensors, sensor fusion, and strict latency for ADAS and autonomy. Enterprise, PC, and personalized edge inference (Priority: 4/5): Post-pandemic PC/enterprise AI growth is being driven by video conferencing, streaming, and personalization use cases using user-specific sensors and data. Data-centric AI and edge MLOps (Priority: 5/5): The discussion highlights a shift from model-centric to data-centric AI, including annotation, filtering, monitoring, drift detection, synthetic data, and possible edge/cloud retraining loops. Transformer and LLM enablement on-device (Priority: 4/5): Supporting LLMs on edge devices requires memory, fixed-point quantization, software tooling, and acceleration for encode/decode workloads, with expectation that major silicon vendors will continue enabling this. Micro-tile inferencing and on-chip parallelism (Priority: 4/5): Qualcomm’s micro-tile inferencing breaks graphs into smaller tiles and dispatches them across specialized cores to reduce latency and power versus running whole graphs serially.

Key Arguments: Edge AI is moving from mostly convolution-heavy vision workloads to broader modalities such as text, linguistics, commerce, recommendations, and generative content. Hardware must be designed years ahead around workload KPIs, especially latency, performance, QoS, power efficiency, memory footprint, and bandwidth. Fixed-point execution is preferred because it can deliver higher performance, better power efficiency, and smaller memory footprint than floating point. Software optimization matters as much as hardware, including quantization, runtime, libraries, compiler support, encryption, and preemption for concurrent workloads. Mobile AI is primarily camera-driven and battery-constrained, while automotive AI is sensor-fusion-heavy, latency-sensitive, and far more compute intensive. Automotive and enterprise use cases are pushing AI requirements to scale faster than initially expected, especially with ADAS, autonomy, video conferencing, and streaming. Edge AI will increasingly require personalized inference based on user-specific data from sensors, cameras, audio, and text. MLOps on the edge must include data collection, labeling, monitoring, drift detection, and possibly federated or hybrid learning so models can improve without sending sensitive data off-device. Hybrid cloud-edge execution remains necessary for heavy workloads like translation and LLMs, but edge capabilities will keep expanding as user expectations for responsiveness rise. Micro-tile inferencing enables more parallel, specialized execution of model subgraphs, improving throughput and power efficiency for multiple concurrent models.

Data Points: Experience in AI/ML: 10-15 years - Vinesh describes the length of his career in AI and machine learning. Transformer/LLM parameter scale: 175 billion+ and 200 billion+ - He references large model sizes when discussing memory and edge feasibility, including OpenAI-style models. Meta LLaMA model size: 7-10 billion parameters - Used as an example of a smaller LLM that people are already trying to run locally. Autonomous driving levels: L1 to L5 - He explains automotive autonomy progression and where compute demand escalates. Common mobile input tensor sizes: 512x512, 225x225 - Compared with automotive camera inputs, mobile tensors are described as relatively small. Automotive camera input sizes: 1, 3, 5, 8 megapixels - He cites higher-resolution tensors required for broader field of view and stricter decisions in cars. AI stack announcement timing: A couple of months ago last year - Reference to Qualcomm AI Stack being newly announced recently relative to the interview. Generative AI startup activity: 150+ startups - He mentions this number in relation to startups building on OpenAI APIs.

Pivotal Quotes: "what you know today is obsolete tomorrow" — Vinesh Sukumar: He explains why working in AI requires continuous learning because the field changes rapidly. "we want to move towards fixed point as we believe it gives you the highest performance, quite leadership-class power efficiency, and also a very small memory footprint" — Vinesh Sukumar: He describes Qualcomm’s hardware/data-type strategy for efficient edge inference. "how can I collect more data about the user to drive personalization?" — Vinesh Sukumar: He frames the future of edge AI as increasingly personalized and sensor-driven.

Implications: Edge AI is moving toward personalized, multimodal, and generative experiences that require new hardware/software co-design, stronger MLOps, and hybrid deployment. Companies that optimize for efficiency, memory, and data quality will be best positioned as on-device AI expands.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast