The TWIML AI Podcast
The TWIML AI Podcast

AI Trends 2024: Reinforcement Learning in the Age of LLMs with Kamyar Azizzadenesheli - #670

Today we’re joined by Kamyar Azizzadenesheli, a staff researcher at Nvidia, to continue our AI Trends 2024 series. In our conversation, Kamyar updates us on the latest developments in reinforcement learning (RL), and how the RL community is taking advantage of the abstract reasoning abilities of lar

Featured Speakers

Kemyar Aziza Denishelli Guest

Topics Discussed

Episode Summary

Executive Summary: The episode argues that 2023 marked a turning point for reinforcement learning as large language models and generative AI began supplying world knowledge, abstraction, and instruction to RL agents. Kemyar Aziza Denishelli says this enables more tractable task decomposition, interactive robotics, reward design, and better risk-aware deployment, while also pushing RL theory toward new formulations that account for LLM-guided exploration and reasoning.

Main Topics: LLMs as a new abstraction layer for RL (Priority: 5/5): The conversation centers on how LLMs and generative AI provide world knowledge, imagination, and task decomposition that RL agents previously lacked, making hard tasks feasible without starting from scratch. Theory of reinforcement learning must be reformulated (Priority: 5/5): Kemyar argues that classical RL foundations should be rethought in light of access to LLMs, especially around hierarchy, abstraction, and exploration-exploitation under richer priors. Practical robotics and interactive agents (Priority: 5/5): The discussion highlights robots that can be instructed in language, use LLMs to break down tasks, and benefit from curriculum learning and human-in-the-loop guidance. Voyager and world-model approaches (Priority: 4/5): Examples like Voyager and world-model-based methods illustrate how code generation, curriculum learning, and imagined goal states can replace brute-force trial and error. Risk-aware RL in production domains (Priority: 4/5): The episode emphasizes increasing use of RL in finance, healthcare, insurance, drones, and control, where optimizing expected reward alone is insufficient and risk constraints matter. Control, drones, and industrial deployment (Priority: 4/5): Kemyar explains that RL is already delivering gains in control systems such as wing stabilization, IMU localization, and drone navigation, often outperforming traditional approximations. Future of RL, compute, and generalization (Priority: 4/5): The future is framed as domain-specific progress now, with broader unification later when compute, data, and theory catch up; the field may shift from game-solving to practical systems and safer AI.

Key Arguments: LLMs supply global world knowledge and abstraction that let RL agents begin from higher-level priors instead of random exploration. LLMs can be used not only to instruct agents in natural language but also to design rewards, generate code, and break tasks into solvable subtasks. Classical RL theory assumed little or no prior knowledge; access to LLMs changes the optimal formulation of exploration, exploitation, and hierarchical planning. Robotics is moving toward interactive, language-assisted training, where agents can be guided, corrected, and evaluated using generative models. Risk-sensitive RL is increasingly necessary in finance, healthcare, insurance, and other high-stakes settings because average reward can hide harmful distributions of outcomes. Industrial control problems often benefit from RL because domain-specific algorithms can outperform hand-tuned approximations, especially when the dynamics are complex or turbulent. A major future research direction is to build LLM-aware RL algorithms that use information gathering and tree-like reasoning rather than purely reflexive policies. The field is becoming more domain-specific in the short term, but these specialized advances may later merge into broader foundations, as has happened in other areas of machine learning.

Data Points: Historical span of classical RL development: 40–50 years - Kemyar refers to classical reinforcement learning topics developed over decades. Approximate performance gain in control: 30–40% better - He says RL can outperform traditional methods for some wing/control problems by this margin. Wind tolerance improvement for drones: 20 m/s vs 5 m/s - He claims RL-enabled drones can handle stronger turbulent wind than before. Training time for drone adaptation: 5 minutes - He describes rapid online adaptation for drones in wind conditions. Field growth in RL papers at NeurIPS: Quadrupled - He says RL submissions at NeurIPS quadrupled in 2016–2017. Expected compute improvement needed: A factor of 10 - He suggests that even 10x better compute would unlock major progress. Suggested societal impact example: 5% wealth increase - Used to illustrate that average gains can hide unequal or harmful distributional outcomes.

Pivotal Quotes: "we now can use these tools as a means of knowledge abstraction" — Kemyar Aziza Denishelli: Explaining why LLMs matter for reinforcement learning research in 2023 and beyond. "The problem is going to be solved in two steps" — Kemyar Aziza Denishelli: Describing how LLMs break complex tasks like making pasta into subgoals. "if you want to do reasoning and if you want to train a model which, given the input state, tells you exactly what is the next move, training such a model is going to be extremely expensive" — Kemyar Aziza Denishelli: Reflecting on AlphaGo-style lessons for future LLM reasoning systems.

Implications: RL is moving from brute-force exploration toward language-guided, risk-aware, domain-specific systems. For industry, this promises better robotics and control. For researchers, it raises open theory questions about how LLMs change optimal RL.

🔓 Sign Up for Unlimited Episode Search

About The TWIML AI Podcast

View all episodes from The TWIML AI Podcast