Episode Summary
Executive Summary: Nathan Lambert explains AI2’s Tulu3 project, which aims to match frontier post-training performance using open methods while releasing data and techniques publicly. The conversation covers SFT, DPO, and verifiable-RL, why data quality dominates algorithm choice, how LLMs can substitute for humans in preference labeling, and why open-source labs may still remain competitive despite closed-model opacity and the rise of o1-style reasoning.
Main Topics: AI2, openness, and the Tulu3 mission (Priority: 5/5): Lambert describes the Allen Institute for AI as an open, hybrid academic-industrial lab funded by Paul Allen’s estate, focused on building credible open language models and using them to make a case for open AI development. Post-training stack: SFT, DPO, and RL (Priority: 5/5): The core technical discussion explains the three-stage post-training pipeline, how each stage changes model behavior, and why the team sees value in combining them rather than obsessing over preference tuning alone. Data quality beats algorithmic novelty (Priority: 5/5): A repeated theme is that careful data curation, decontamination, and targeted dataset design produce bigger gains than algorithmic tweaks, especially in post-training. LLM-generated preference data and on-policy training (Priority: 4/5): Lambert argues that LLMs can replace humans for many mechanical labeling tasks, and that on-policy preference data from current model generations is better than generic off-the-shelf preference sets. Verifiable-reward RL and emergent reasoning (Priority: 5/5): The team adds a final RL stage using objective verifiers for math and formatting tasks, which can recover or improve performance and sometimes induces o1-like self-checking behavior. Evaluation, contamination, and benchmarking (Priority: 4/5): The conversation emphasizes how fragile model evaluations are, how easy contamination is, and why decontamination and careful eval design are now essential parts of serious model development. Open-source versus closed-model future (Priority: 4/5): Lambert suggests the open community can still adapt, but the loss of frontier reasoning traces and the increasing sophistication of closed labs raise uncertainty about how open post-training will evolve.
Key Arguments: The biggest leverage in post-training comes from better data pipelines and evaluation, not from endlessly tuning preference algorithms. Open labs can partially substitute LLMs for humans in preference labeling because many labeling tasks are mechanical, though this introduces bias. On-policy preference data is better because it matches the current model distribution more closely than generic datasets. Verifiable-reward RL is promising because objective checks for math, formatting, and other constraints can provide scalable supervision. SFT does most of the heavy lifting for general capability gains, while DPO and RL usually add the final percentage points. Data contamination is a major hidden problem; even small overlaps can meaningfully distort leaderboard results. Model character is important but hard to evaluate, which is one reason open models still lag closed models like Claude in consistency. The open ecosystem likely can keep up, but only if it develops new infrastructure and new ways to bootstrap reasoning-like data without access to proprietary traces.
Data Points: Tulu iteration count: Tulu 3 / Tulu3 - The latest AI2 open post-training project discussed in the episode. Team size: 10-15 active contributors - Lambert says the project was driven by a relatively small team despite the scale of the work. Base-model scale: 8B and 70B - The project trained and validated models at both sizes to ensure recipe transfer. Training runs: ~1,000 8B models - Used to search the post-training space and optimize the recipe. SFT dataset size: ~1,000,000 prompts - Final supervised fine-tuning mix used for the general-purpose instruction model. SFT throughput: 32 H100s for about 1 day - Approximate time to train an 8B model on the SFT stage. DPO data size: a couple hundred thousand prompts - Order-of-magnitude estimate for preference-tuning data volume. DPO runtime: 6-12 hours - Typical runtime on 16-32 GPUs depending on scale and implementation details. Compute cost mention: about $1,500 for SFT compute ballpark - A rough back-of-the-envelope estimate discussed for one SFT training cycle. API spend: over $50,000 - Lambert estimates LLM-as-judge/API costs for the project. Math mix share: 30%+ - He says math made up a large portion of the SFT mix as the team targeted math benchmarks. GSM8K uplift from verifiable RL: about 15 points - He cites prior older models gaining substantially from RL on verifiable outputs. Checkpoint count on leaderboard: 1,000+ models - The team evaluated many variants while searching for the best recipe. Contamination threshold: >2% exact prompt match - Training sets with more than 2% exact prompt overlap with eval prompts were treated as contaminated and removed. Quantization of Llama 3.1 performance: 8B and 70B matching/beat milestones in weeks - Lambert describes the project passing Llama 3 instruct and then Llama 3.1 instruct shortly after starting. Preference-tuning contribution: last 10% - His rough rule of thumb is that SFT delivers about 90% of the final performance, with DPO and RL making up the rest. RL KL scale: ~1 for basic math RL - He contrasts this with much larger KL spends in chat-oriented PPO setups. Human vs machine labeling: high-noise/low-bias vs low-noise/high-bias - He quotes John Schulman’s framing of human and LLM preference data. OpenAI/Google score trend: “skyrocketing” / “hockey stick” - Used to describe the renewed steep gains in frontier benchmarks and Chatbot Arena.
Pivotal Quotes: "It’s probably not worth the effort to spend all your time on preference tuning when you can just be making better data and better pipelines." — Nathan Lambert: He summarizes the project’s philosophy and why Tulu3 emphasizes data and process over algorithmic obsession. "Human preference data is high noise but low bias, and LLM preference data is low noise but high bias." — Nathan Lambert: He explains why AI2 uses LLMs for many annotation tasks while acknowledging the tradeoffs. "The big thing is, how do you develop character? Character is something that you don’t have a valuation for in your models." — Nathan Lambert: He describes one of the hardest-to-measure aspects of model quality and why open models may lag closed ones in consistency.
Implications: Open model teams can still compete, but only by mastering data curation, decontamination, on-policy labeling, and verifier-based RL. The next frontier likely centers on reasoning-style training, evaluation integrity, and new open infrastructure—not just bigger models.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co