Catalyst with Shayle Kann
Catalyst with Shayle Kann

Can AI revolutionize materials discovery?

AI is working its way across climate tech, helping companies discover giant lodes of ore, catch battery defects, and monitor energy infrastructure. Could it help us find revolutionary new materials, too? Turns out, it’s complicated. In this episode, Shayle talks to Ekin Dogus Cubuk, or Dogus, a rese

Featured Speakers

Doge Chubuk Guest

Topics Discussed

Episode Summary

Executive Summary: The episode examines whether AI can become a climate-tech “killer app” through materials discovery. DeepMind researcher Doge Chubuk argues AI and simulation can accelerate optimization of known materials, but true novelty remains hard because models are constrained by limited, noisy data and weak extrapolation beyond known distributions. The most promising near-term gains appear in batteries, some catalysts, and improving DFT, not yet in paradigm-shifting discoveries like room-temperature superconductors.

Main Topics: AI’s role in climate tech beyond energy demand (Priority: 5/5): Shail frames two AI-climate intersections: AI’s rising power demand from data centers, and the less-discussed opportunity to use AI itself for climate solutions, especially materials discovery. How materials discovery worked before AI (Priority: 5/5): Chubuk explains that historically materials science relied heavily on trial and error, serendipity, and partial physical intuition, citing examples from metallurgy, transistors, and batteries. Limits of AI for truly novel discovery (Priority: 5/5): The discussion emphasizes that ML and simulation are best at interpolation and optimization within known distributions, while discovering wholly new materials requires extrapolation beyond training data where current models struggle. Where AI can help today: optimization and screening (Priority: 4/5): AI may speed up incremental improvements in known classes of materials, such as refining superconductors, batteries, or other systems where the design space is partially understood. Data scarcity and noisy labels in materials science (Priority: 5/5): Unlike internet-scale LLM training, materials data is sparse, incomplete, and often noisy; only a small subset of materials have well-characterized properties, limiting model performance. Promising and less promising application areas (Priority: 4/5): Chubuk suggests batteries are a better fit for current methods than catalysis or superconductivity, while optical/electronic properties face constraints from the accuracy of DFT. DeepMind’s current work and future goals (Priority: 4/5): Chubuk describes GNOME, which found stable crystals at 0 Kelvin, and says the team aims to extend to finite-temperature stability and improve DFT with machine learning.

Key Arguments: Materials discovery has long depended on serendipity and incomplete intuition, not systematic exhaustive search, which is why AI is attractive. Current AI and simulation methods are good at optimizing within a known materials family, but poor at discovering completely new paradigms. The farther a prediction is from known data, the less reliable the approximation becomes; this mirrors both machine learning and scientific modeling. Materials science lacks the scale and label quality that made LLMs and AlphaFold successful, making a breakthrough model harder to achieve. Experimental validation remains essential because computational models cannot yet reliably replace real-world measurements. Batteries are a comparatively strong near-term target because stability and ion transport are more directly modelable than messy surface or many-body phenomena. AlphaFold-like watershed moments are possible only if the field develops a high-quality, objective benchmark dataset and consistent experimental labels. Improving DFT with machine learning may yield practical gains even if AI does not immediately discover radically new materials.

Data Points: Inorganic crystal structure database size: more than 200,000 crystals - Chubuk cites ICSD as a rough upper bound of known inorganic crystal data available for materials modeling. Known property-labeled materials: about 1,000 to 2,000 data points - He notes only a small subset of materials have well-characterized properties like band gap or conductivity. Computational training points in earlier work: several million - He says DeepMind’s earlier paper used millions of simulation-derived training examples. Larger computational datasets now: around 50 million points - He says multiple groups are now pushing simulation datasets to tens of millions of examples. Zero-Kelvin stability training data: at most 48,000 total predictions - For the GNOME project, the relevant stability-prediction dataset was limited compared with internet-scale AI. Zero-Kelvin stability data from computation: about 28,000 - Chubuk says approximately 28,000 data points came from computation in the GNOME dataset. Zero-Kelvin stability data from experiments: about 28,000 - He says roughly 28,000 came from prior experiments, underscoring the field’s small scale. Episode ad metric: 2.5 million customer devices - Energy Hub ad claims its VPP aggregates customer devices into dispatchable grid capacity. Episode ad metric: 3.4 gigawatts - Energy Hub says its device fleet equals 3.4 GW of dispatchable capacity. Episode ad metric: more than three nuclear reactors - The ad compares 3.4 GW of flexible capacity to three-plus nuclear reactors. Episode ad metric: more than 170 utilities - Energy Hub ad says utilities are using its VPP platform during peak season. Episode ad metric: over 25 years - Bloom Energy ad references a 25+ year track record with hospitals, universities, and utilities.

Pivotal Quotes: "The farther we get from the sphere, the less good our approximations will be." — Doge Chubuk: Explaining why AI and simulation are strongest near known materials and weaker for truly novel discoveries. "AI and simulations can be useful here because a lot of the discoveries have just been randomly trying kind of relevant materials." — Doge Chubuk: Describing why materials discovery is a good candidate for machine-learning acceleration. "There is a certain amount that we know as humans, and maybe we can use computation to predict a bit outside of that circle sphere." — Doge Chubuk: A philosophical framing for the limits of extrapolation in scientific modeling.

Implications: AI is likely to accelerate materials optimization, not instantly invent breakthrough materials. Near-term value is in better screening, faster iteration, and improved DFT; major climate wins will still require experiments, better datasets, and domain-specific benchmarks.

🔓 Sign Up for Unlimited Episode Search

About Catalyst with Shayle Kann

View all episodes from Catalyst with Shayle Kann