The TWIML AI Podcast
The TWIML AI Podcast

High-Efficiency Diffusion Models for On-Device Image Generation and Editing with Hung Bui - #753

In this episode, Hung Bui, Technology Vice President at Qualcomm, joins us to explore the latest high-efficiency techniques for running generative AI, particularly diffusion models, on-device. We dive deep into the technical challenges of deploying these models, which are powerful but computationall

Featured Speakers

Hung Bui Guest

Topics Discussed

Episode Summary

Executive Summary: Hung Bui traces his path from early AI curiosity and work on Siri at SRI to founding VINAI in Vietnam, where resource limits pushed a research strategy centered on efficiency. The lab built smaller but stronger Vietnamese language models, one-step diffusion image generation (SwiftBrush), one-step image editing (SwiftEdit), and is now pushing on-device multimodal agents and inference-time scaling, all reinforced by Qualcomm’s acquisition and an expanded residency program.

Main Topics: Hung Bui’s AI career arc (Priority: 5/5): Bui describes how curiosity about the Turing test led him from a PhD in multi-agent systems to roles at SRI, Nuance, Adobe, and DeepMind, culminating in building an AI lab in Vietnam. Building VINAI in Vietnam (Priority: 5/5): He explains the challenge and opportunity of starting a world-class AI lab in Vietnam, including the strategy of recruiting experienced returnees and training local talent through an AI residency program. Efficiency as a research strategy (Priority: 5/5): Limited compute in Vietnam forced the lab to focus on making models smaller and more efficient rather than simply scaling up, shaping much of its work in generative AI. Vietnamese language model scaling (Priority: 5/5): The lab trained a Vietnamese-only model first at 7B parameters, then under 4B, discovering that careful training improvements let the smaller model outperform the larger one. One-step diffusion for image generation and editing (Priority: 5/5): Bui details SwiftBrush and SwiftEdit, which distill multi-step diffusion into one-step generation/editing, dramatically reducing latency while preserving strong quality. On-device agents and multimodal AI (Priority: 4/5): He argues that personalized agents need access to private device data, which makes on-device multimodal processing, memory, and retrieval core research priorities. Inference-time scaling and future optimization (Priority: 4/5): Bui frames test-time scaling as both a challenge and an opportunity for constrained devices, since it can make smaller models competitive with much larger ones on reasoning tasks.

Key Arguments: Resource limits can drive better research directions: because VINAI lacked massive compute, it focused on efficiency, smaller models, and better training recipes. A 7B Vietnamese-only model was enough to produce useful Vietnamese chat-style behavior, showing that frontier-like behavior can emerge at much smaller scale than global flagship models. Reducing model size below 4B parameters and changing the training approach yielded better performance than the 7B model, validating efficiency-first optimization. One-step diffusion is feasible through distillation of the multi-step denoising process, enabling fast image generation and editing with near-teacher quality. For image generation, the bottleneck is often repeated denoising steps rather than model size, so reducing steps can yield major latency gains. On-device agents are important because personalization requires access to private local data, and processing that data on-device improves privacy. Test-time scaling is a double-edged sword for mobile but can make small models outperform much larger ones on certain tasks, creating a strong research opportunity. The residency program solved talent constraints by blending experienced hires with highly trainable local researchers, and it has become a major pipeline for AI talent in Vietnam.

Data Points: PhD timing: almost 30 years ago - Hung Bui says his PhD in multi-agent systems was completed almost 30 years prior to the interview. First lab launch: 2019 - He moved back to Vietnam in 2019 to help set up VINAI, the first AI research lab in the country. Initial Vietnamese LLM size: 7 billion parameters - VINAI trained an early Vietnamese-only generative model at 7B parameters. Reduced model size: less than 4 billion parameters - The lab later reduced the model size below 4B and improved performance through training changes. ChatGPT-3 scale reference: close to 200 billion parameters - Bui cites ChatGPT-3 as a scale reference that was far beyond VINAI’s available compute. Diffusion inference steps: 50 to 100 time steps - He notes standard denoising diffusion image generation often requires dozens to around 100 repeated steps. Latency improvement target: almost two order magnitudes - Bui estimates one-step generation could improve latency by nearly 100x versus multi-step diffusion. Image editing runtime: a fraction of a second / a quarter of a second - SwiftEdit is described as running on a single GPU in real time, around a quarter of a second per edit. Total lab size: about 90 people - Current total staff in the lab is around 90 researchers and engineers. Residents in lab: between 40 and 50 - At any time, roughly 40-50 residents are in the lab. Residency duration: 2 years - Residents stay for two full years, long enough to contribute to research and engineering. Alumni count: close to 100 - Nearly 100 people have gone through the residency program over time.

Pivotal Quotes: "we noticed that this model, less than 4 billion parameters, actually perform even better than 7 billion" — Hung Bui: On the surprise result that a smaller Vietnamese model outperformed the larger 7B version after training improvements. "the goal is the same, right? How you be able to get text generation, and also here's how you'll be able to generate an image, but in a most efficient way" — Hung Bui: Summarizing the lab’s consistent efficiency-first philosophy across text and image generation. "we want this agent or assistants need to have access to our private information. Otherwise, how can it actually be personalized?" — Hung Bui: Explaining why on-device processing is central to building personalized agents.

Implications: The interview shows that frontier-like capabilities do not require maximal scale if training and distillation are optimized. It also suggests a future where privacy-preserving, multimodal, on-device AI and test-time reasoning become key battlegrounds for mobile and edge AI.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast