Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

RLHF 201 - with Nathan Lambert of AI2 and Interconnects

In 2023 we did a few Fundamentals episodes covering Benchmarks 101, Datasets 101, FlashAttention, and Transformers Math, and it turns out those were some of your evergreen favorites! So we are experimenting with more educational/survey content in the mix alongside our regular founder and event cover

Featured Speakers

Latent.Space HostNathan Lambert Guest

Topics Discussed

Episode Summary

Executive Summary: Nathan Lambert gives a deep, practical tour of RLHF: its intellectual roots, why preference data matters more than flashy RL theory, how instruction tuning, reward models, PPO, and DPO fit together, and why evaluation remains messy. He argues modern alignment is mostly a data/infrastructure problem, with RL serving as one tool among many, and highlights emerging directions like synthetic feedback, constitutional AI, and open-source DPO models.

Main Topics: RLHF’s intellectual history and assumptions (Priority: 5/5): Lambert traces RLHF from economics, philosophy, and classical RL to modern LLM alignment, emphasizing that preference modeling rests on deep assumptions about whether human preferences can be measured and aggregated. Reinforcement learning vs. language-model adaptation (Priority: 5/5): He explains that RLHF is only loosely analogous to classical RL: language models lack a true environment/state loop, so the setup is closer to bandits or preference optimization than full sequential RL. Instruction tuning as the practical foundation (Priority: 4/5): Instruction tuning is presented as the most useful and accessible step for most teams, often delivering the biggest immediate gains before any RLHF stage is attempted. Preference data collection and reward modeling (Priority: 5/5): The conversation details pairwise comparisons, Likert-style interfaces, metadata labeling, and Bradley-Terry-style reward models, stressing that data quality and annotation design are central to success. PPO, KL constraints, and why RLHF is hard to scale (Priority: 4/5): Lambert describes the standard RLHF objective, the role of KL regularization, and the complexity of PPO-based training, arguing that the engineering burden is substantial and often underappreciated. Emerging alternatives: rejection sampling, constitutional AI, DPO (Priority: 5/5): The discussion covers cheaper or simpler alternatives to PPO, including rejection sampling, AI-generated feedback, constitutional AI, and direct preference optimization, which is gaining traction in open source. Evaluation, benchmarks, and real-world model quality (Priority: 4/5): He argues that benchmarks are useful but insufficient, and that direct interaction with models and arena-style evaluations often reveal more than static leaderboards.

Key Arguments: RLHF is built on a chain of assumptions from economics and philosophy, not just machine learning; the core question is whether human preferences can be modeled and aggregated at all. In language models, RLHF is not “real RL” in the classical sense because there is no clean environment-state-action loop; it is closer to preference-based optimization or bandits. Instruction tuning is still the most practical and broadly useful technique for most teams; many organizations do not need the complexity of full RLHF. Preference data quality matters more than raw quantity, but collecting it is expensive, noisy, and often inconsistent across annotators. Pairwise preference modeling works better than scalar ratings because humans are better at comparative judgments than absolute scoring. The KL constraint is crucial because RLHF data is sparse and models can overfit badly without a guardrail against drifting too far from the base model. Open-source and academic teams are increasingly adopting DPO because it is simpler to implement and often performs well enough, even if PPO may still have a higher ceiling in some settings. Evaluation remains unresolved: benchmarks can be gamed, and the best way to judge a model is still to talk to it and test it in realistic use cases. Synthetic feedback, especially from GPT-4, is becoming attractive because it is far cheaper than human labeling and can approximate human preference signals reasonably well. Constitutional AI and related methods are less about explicit moral encoding than about using AI-generated critiques and sampled principles to steer model behavior.

Data Points: Trailhead distance from home: 1.2 miles - Lambert says he has a trailhead very close to his house in the Bay Area, supporting his ultra-endurance hobby. RLHF data agreement: 65%–75% - He cites typical agreement/accuracy levels for reward models against preference data, showing that perfect alignment is not expected. Human agreement on preferences: about 60%–70% - He notes that humans themselves often only agree at this level when labeling preferences. GPT-4 preference-labeling accuracy: about 80% - He claims GPT-4 can outperform humans on some preference-labeling tasks, making synthetic feedback attractive. Llama 2 preference-data cost estimate: $6M–$8M - He revises an earlier estimate and says this is a safer ballpark for the preference-data portion of Llama 2. Earlier rough preference-data estimate: $20M–$30M - He says his earlier estimate was off by a factor of four due to converting prompts to total data points incorrectly. Llama 2 preference data volume: 1.5 million data points - He references the Llama 2 paper’s reported scale of preference data. Prompt count equivalent: about 400,000 prompts - He explains that the 1.5 million data points may correspond to roughly 400,000 prompts in a multi-turn setting. Vicuña budget claim: under $100 - He cites the open-source instruction-tuning effort as an example of how far a small budget can go. GPT-4 Arena score gap: 40 ELO points - He notes GPT-4 March 14 is about 40 ELO points above GPT-4 June 13 in LMSys Chat Arena.

Pivotal Quotes: "RLHF, while directly building on tools from RL and language models, is really implicitly impacted and built on theories and philosophies spanning tons of human history." — Nathan Lambert: He explains the intellectual lineage of RLHF beyond standard ML. "I think in the long run, it will still settle out where RL will still be a field that people work on just because of these kind of fundamental things that I talked about that it's just viewing the whole problem formulation different than predicting text." — Nathan Lambert: He distinguishes classical RL from LLM training and argues the field will remain distinct. "I think it's getting good data." — Nathan Lambert: He answers what matters most for the next generation of model builders: data quality over theoretical elegance.

Implications: For builders, the message is clear: prioritize high-quality data, strong evaluation, and simple methods before complex RL. For the industry, synthetic feedback, DPO, and constitutional AI will likely expand, while human preference data remains costly and strategically valuable.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast