The TWIML AI Podcast
The TWIML AI Podcast

Gauge Equivariant CNNs, Generative Models, and the Future of AI with Max Welling - TWiML Talk #267

Today we’re joined by Max Welling, research chair in machine learning at the University of Amsterdam, and VP of Technologies at Qualcomm, to discuss: • Max’s research at Qualcomm AI Research and the University of Amsterdam, including his work on Bayesian deep learning, Graph CNNs and Gauge Equivaria

Featured Speakers

Max Welling Guest

Topics Discussed

Episode Summary

Executive Summary: Max Welling traces his path from theoretical physics to computer vision and machine learning, then discusses founding Cypher, its acquisition by Qualcomm, and his current work bridging academia and industry. The conversation centers on efficient compute, Bayesian deep learning, symmetry- and gauge-equivariant neural networks, and a hybrid view of AI that combines data-driven models with generative/causal reasoning for better generalization and practical deployment.

Main Topics: Career path from physics to machine learning (Priority: 5/5): Welling explains how he moved from theoretical physics to computer vision and then machine learning, finding a field with broader real-world impact and rapid growth. Founding Cypher and Qualcomm acquisition (Priority: 5/5): He describes how a bank competition led to Cypher’s creation, its consultancy-driven growth, an active-learning defect-detection product, and eventual acquisition by Qualcomm. Compute efficiency, compression, and hardware-aware AI (Priority: 5/5): A major theme is the importance of compute as models scale, including neural network compression, quantization, and optimizing algorithms for mobile and specialized hardware. Bayesian deep learning and uncertainty (Priority: 4/5): Welling discusses using Bayesian methods to compress neural networks by large factors while also enabling uncertainty estimates over predictions. Symmetry, capsules, and gauge-equivariant neural networks (Priority: 5/5): He connects group symmetries, capsules, and gauge theory to modern deep learning on non-Euclidean domains such as manifolds, meshes, and curved surfaces. Data-driven vs generative/causal AI (Priority: 5/5): The interview explores Rich Sutton’s data-centric argument versus generative modeling and causal understanding, proposing a middle ground depending on task and data availability. Hybrid systems and switching between models and rules (Priority: 4/5): Welling argues that future AI systems will likely combine learned models, generative reasoning, and rule-based safeguards, especially in safety-critical domains like autonomous driving.

Key Arguments: Welling’s shift from physics was motivated by wanting a field with more immediate and dynamic real-world impact, which he found in machine learning. Cypher’s origin was opportunistic: a small model built by a student on a laptop outperformed much larger consulting firms in a bank ad-click competition. Consultancy work across finance, manufacturing, and retail gave the team broad practical experience and helped make Cypher attractive for acquisition. Compute is not just an implementation detail; it is central to deep learning progress, and future gains require new compute paradigms that reduce memory movement and energy use. Neural networks are often heavily over-parameterized; Bayesian deep learning helped show that models can be compressed dramatically without losing accuracy. Symmetry is a powerful inductive bias: incorporating translational, rotational, and local gauge symmetries improves neural networks and generalizes naturally to manifolds. Gauge-equivariant methods are especially useful when learning on non-flat domains such as the Earth, hearts, meshes, or other curved structures. Data-driven models excel when there is lots of training data in a narrow domain, but they struggle to generalize to rare corner cases and new environments. Generative/causal models are better suited to broad generalization and reasoning from first principles when data is sparse or the environment changes. Future AI systems will likely need a hybrid strategy: use rules or generative models in unfamiliar or safety-critical cases, and switch to learned models when enough data exists. Hardware and software are becoming increasingly co-designed; optimizing algorithms for specific chips and compute-in-memory approaches will matter more over time.

Data Points: Cypher acquisition timing: 2017 - Welling says Qualcomm acquired Cypher in 2017. Joint appointment split: About half academia, half Qualcomm - He describes balancing his time between the University of Amsterdam and Qualcomm. Neural network compression factor: 100x - He says Bayesian deep learning enabled compression of a neural network by a factor of 100 without losing accuracy. Model size: Hundreds of millions to billions of parameters - He cites the scale of modern neural networks as a compute and memory challenge. Hardware rollout: A billion phones - He notes Qualcomm’s ability to ship technology at massive scale into mobile devices. Temporal horizon at Qualcomm: 1 to 5 years - He contrasts Qualcomm’s applied research timeframe with academia’s longer horizons. Academic horizon: Potentially 100 years out - He describes academia as free to pursue very long-term research. Bank competition: Single GPU laptop model beat big companies - Cypher began after a student’s laptop implementation outperformed large consulting teams in an ad-click prediction competition. Deep learning domain example: Speech and computer vision - He uses these as examples where data-rich approaches outperformed hand-built/mechanistic models.

Pivotal Quotes: "I basically beat all the big companies." — Max Welling: Describing the bank competition that sparked the founding of Cypher. "We need ways to generalize the lessons we learn in one context into sort of a completely different context." — Max Welling: On the limits of narrow data-driven models and the need for broader reasoning. "If I don't know much about a domain, if I get thrown into a new situation, I need to rely on my generative models of the world." — Max Welling: Explaining his preferred hybrid approach for general AI and unfamiliar environments.

Implications: AI’s near future likely depends on combining efficient hardware, compressed models, symmetry-aware architectures, and hybrid reasoning systems. Narrow tasks will keep benefiting from scale and data, but robust real-world AI will need uncertainty, causal structure, and fallback rules.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast