The Cognitive Revolution
The Cognitive Revolution

Nested Learning: Ali Behrouz on the Quest for Continual Learning & Illusion of AI Architectures

Ali Behrouz, grad student at Cornell and Google researcher, discusses his potentially transformative work on new architectures for continual learning in AI. His paper "Nested Learning," praised by Jeff Dean as a possible paradigm shift, enables models to adapt to new context while preservi

Featured Speakers

Nathan Labenz and Erik Torenberg HostAli Beyrouz Guest

Topics Discussed

Episode Summary

Executive Summary: Ali Beyrouz argues that today’s AI systems lack true continual learning, stable memory consolidation, and flexible abstraction. He proposes nested learning and sleep-like offline updates as a path to architectures that adapt in real time, preserve long-term knowledge, and improve on hard tasks while raising major privacy, alignment, and ecosystem risks.

Main Topics: Continual learning as the central missing capability (Priority: 5/5): Beyrouz frames current LLMs as powerful but fundamentally limited because they cannot continuously learn new knowledge without catastrophic forgetting or inefficient full-model updates. Nested learning: multiple update frequencies (Priority: 5/5): Nested learning replaces a single update regime with several memory levels that update at different frequencies, mirroring human-like short-, medium-, and long-term memory. HOPE architecture and self-modifying associative memory (Priority: 5/5): The HOPE family extends transformer-like systems by using multiple MLP blocks or Titan-like modules with different update rates, including a self-modifying variant where the model learns its own update rule/value generation. Sleep phase: memory consolidation and dreaming (Priority: 4/5): The Language Models Need Sleep paper adds an offline phase where models distill fast-learned knowledge into slower layers and generate synthetic data to consolidate and abstract from recent experience. Empirical performance and micro-skills (Priority: 4/5): The new architectures are presented as competitive with transformers on standard metrics while outperforming them on harder continual-learning-style tasks such as multi-language in-context translation and noisy recall. Optimizers as nested learning systems (Priority: 3/5): Beyrouz extends the same conceptual framework to optimization, arguing that architecture and optimizer are both nested learning problems and introducing an optimizer variant that outperforms Adam and Muon in some settings. Safety, privacy, and ecosystem implications (Priority: 5/5): Continual learning could personalize models and create a diverse AI ecology, but also introduces serious risks of privacy leakage, value drift, adversarial manipulation, and winner-take-all dynamics.

Key Arguments: Current LLMs are excellent at next-token prediction but weak at continual adaptation, durable memory, and stable identity across interactions. A model that truly learns continuously should have active and offline phases; sleep-like processing can consolidate knowledge even when no user input is present. Different update frequencies let some parts of a model remain stable while others adapt quickly, closely matching how human memory operates across time scales. Nested learning treats many machine learning components as associative memory systems, making deep learning a special case of a broader framework. Knowledge transfer between fast and slow modules is essential; distillation helps preserve useful information while forcing higher-level abstraction. Self-modifying memory mechanisms are more expressive than fixed update rules because the model can learn how to update itself based on context. Attention remains important because it behaves like an extremely fast, near-infinite-frequency memory module and is well-suited for certain tasks. Standard perplexity benchmarks are useful for showing compatibility, but they do not fully capture the core goals of continual learning or long-context adaptation. Continual learning may improve personalization and long-context performance, but it also creates new concerns around privacy, alignment drift, and adversarial influence. A healthier AI future may involve an ecology of diverse systems rather than a single monolithic model that absorbs all learning and dominates the market.

Data Points: Nested Learning development time: more than 1 year and a half - Beyrouz says the nested learning framework took a long time to formalize and develop with coauthors. Model scale in paper: 760 million parameters - One of the HOPE experiments scaled models to this size. Training tokens: 30 billion tokens - Paired with the 760M-parameter model in the reported experiments. Larger model scale: 1.3 billion parameters - A second HOPE model size used in experiments. Training tokens for larger model: 100 billion tokens - Paired with the 1.3B-parameter model in the reported experiments. Frequency schedule example: 128, then 4×128, then 4×128 - Approximate initial chunk-size/frequency settings used for different update levels. Context length stress test: up to 10 million tokens - Mentioned in the episode intro as a hard task where the new architectures perform strongly. Translation benchmark: multiple previously unseen languages at the same time - A highlighted qualitative capability of the new architectures. Two-language failure case: performance almost collapses - When a transformer is asked to learn and translate two unseen languages in the same context, it struggles badly; multiple HOPE levels recover performance.

Pivotal Quotes: "A true continual learner doesn't have a test and train time." — Ali Beyrouz: Explaining why continual learning should be thought of as a uniform process rather than separate training and evaluation phases. "Everything that we know of somehow is a form of in-context learning." — Ali Beyrouz: Describing the nested learning thesis that many learning mechanisms can be reinterpreted as associative memory operating on context. "It is the responsibility of the knowledge transfer methods to avoid such cases." — Ali Beyrouz: On the privacy and alignment risks of continual learning and the need to filter adversarial or misleading updates.

Implications: If continual learning scales, AI may become far more personalized, capable, and context-aware than today’s chatbots. But it will also need new safeguards for privacy, drift, and manipulation—and may push the field toward diverse specialized systems, not one dominant model.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution