The TWIML AI Podcast
The TWIML AI Podcast

Gen AI at the Edge: Qualcomm AI Research at CVPR 2024 with Fatih Porikli - #688

Today we’re joined by Fatih Porikli, senior director of technology at Qualcomm AI Research. In our conversation, we covered several of the Qualcomm team’s 16 accepted main track and workshop papers at this year’s CVPR conference. The papers span a variety of generative AI and traditional computer vi

Featured Speakers

Fatih Perikli Guest

Topics Discussed

Episode Summary

Executive Summary: Sam Charrington interviews Fatih Perikli of Qualcomm AI Research about Qualcomm’s strong CVPR showing and its push to make generative AI, multimodal reasoning, and traditional vision models efficient enough for edge devices. The discussion covers papers on diffusion acceleration, video grounding, coached fitness feedback, math plot reasoning, speculative decoding, prompt-guided image generation, optical flow, stereo compression, and on-device demos/workshops.

Main Topics: Qualcomm’s CVPR research direction (Priority: 5/5): Fatih frames the team’s 16 accepted papers as part of a broader shift toward generative AI and multimodal systems that must run efficiently on mobile, XR, automotive, robotics, and IoT devices. Clockwork Units: accelerating diffusion U-Nets (Priority: 5/5): A text-to-image diffusion optimization that approximates middle U-Net layers, exploiting their lower sensitivity to improve speed without harming perceived quality, saving over 30% of compute. Multimodal grounding in video and fitness coaching (Priority: 5/5): Look, Remember, and Reason and What to Say and When to Say It both teach video-language models to attend to object identity, position, motion, and action quality through targeted training questions and feedback annotations. New benchmarks for fine-grained visual reasoning (Priority: 4/5): Math Search exposes weaknesses in current multimodal models on plot reasoning, while Fit Coach adds a large-scale benchmark for situated exercise feedback, both intended to push local-detail understanding. Inference and generation efficiency techniques (Priority: 5/5): Speculative decoding for multimodal models and segmentation-free guidance for diffusion both show ways to improve speed or quality via architectural/inference tweaks rather than retraining from scratch. Vision fundamentals: optical flow and stereo compression (Priority: 4/5): Several traditional CV papers extend Qualcomm’s work on optical flow augmentation, self-cleaning flow refinement, and low-latency stereo streaming to improve motion understanding and bandwidth efficiency. On-device demos and workshops (Priority: 3/5): Qualcomm highlights practical demos such as adaptable LoRA style control, mobile VLMs, relighting, and autonomous-driving data augmentation, plus workshops on efficient large vision models and omnidirectional vision.

Key Arguments: Generative AI is reshaping Qualcomm’s research agenda, but the company’s focus is making these models practical for edge deployment across mobile, XR, automotive, robotics, and IoT. Not all layers of a diffusion U-Net contribute equally; the middle layers can be approximated to reduce compute while preserving quality metrics like FID and CLIP. Video-language models often capture global context but miss object-level motion and identity unless training explicitly probes those details with targeted questions. Situated feedback models should go beyond generic correctness and provide actionable coaching, which requires datasets annotated with feedback, not just Q&A. Benchmarks like Math Search reveal a major gap between human and model performance on fine-grained chart/plot reasoning, indicating significant room for progress. Speculative decoding can extend beyond pure language models to multimodal systems, improving token throughput by pairing a draft model with a larger verifier model. Prompt-guided diffusion can be improved by adjusting attention behavior rather than modifying the base model, enabling better localized guidance and higher user preference. Traditional computer vision problems like optical flow and stereo compression remain important because they underpin video quality, compression, and real-time perception on constrained hardware.

Data Points: Accepted papers at CVPR: 16 - Qualcomm AI Research papers discussed in the episode Main conference papers authored by Fatih Perikli: 3 - Among the CVPR papers mentioned Workshop papers authored by Fatih Perikli: 4 - Among the CVPR papers mentioned Compute savings in Clockwork Units: more than 30% - Approximation of middle U-Net layers in diffusion models Fit Coach dataset size: more than 470 hours - Exercise video dataset for situated interactions Fit Coach QA pairs: more than 1.7 million - Large-scale annotated training benchmark Math Search dataset size: more than 200,000 example plots - Benchmark for multi-hop visual reasoning over plots Math Search test/validation questions: more than 2,000 - Evaluation set for plot reasoning benchmark Human accuracy on math plots: more than 70% - Approximate performance cited for humans on Math Search-style tasks Best model accuracy on math plots: around 25% - GPT-4V cited as the best baseline in the discussion Multimodal speculative decoding speedup: up to 2.3x - Acceleration achieved using a 150M-parameter draft model Draft model size: 150 million parameters - Used in speculative decoding for multimodal language models Human preference for segmentation-free guidance: 60% - Human study preference for the modified diffusion approach Preference for original model: 19% - Human study comparison in segmentation-free guidance paper Bitrate savings on Cityscapes: 50% - Low-latency neural stereo streaming results Bitrate savings on KITTI 2012/2015: more than 20% / 15% - Stereo compression benchmark results

Pivotal Quotes: "It says, oh smooth on the way down. You know, it's literally useful information to me." — Fatih Perikli: Explaining why coached feedback is more valuable than generic exercise completion feedback "We want this model look at myself, and I'm doing some squats, and better, and I ask provide feedback to me, right? Ordinary rigor feedback would be: ah, the user you have successfully completed a squat. You know, this is okay, it is not incorrect, it's correct, but it's not useful." — Fatih Perikli: Describing the motivation for the situated interactions fitness model "The gap is actually quite large, you know, quite big. On average, let's say human does 70% human accuracy, let's say around more than 70%. But these models, even the best one, which is GPT-4V, is around 25%." — Fatih Perikli: Summarizing the benchmark gap revealed by Math Search

Implications: The episode shows Qualcomm pushing a practical AI agenda: richer multimodal understanding plus efficiency for edge deployment. The benchmarks and methods point to faster, more useful on-device assistants, better media generation, and clearer evidence that current models still struggle with fine-grained reasoning.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast