Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

The RLVR Revolution — with Nathan Lambert (AI2, Interconnects.ai)

We first had Nathan on to give us his RLHF deep dive when he was joining AI2, and now he’s back to help us catch up on the evolution to RLVR (Reinforcement Learning with Verifiable Rewards), first proposed in his Tulu 3 paper. While RLHF remains foundational, RLVR has emerged as a powerful approach

Featured Speakers

Latent.Space HostNathan Lambert Guest

Topics Discussed

Episode Summary

Executive Summary: Nathan Lambert discussed AI2’s Tulu and RLVR work, arguing that open-model post-training is converging on industry practice through verifiable rewards, tool use, and multi-stage reasoning. He emphasized that the next frontier is less about new algorithms and more about data, curricula, planning, calibration, and environment design. He also contrasted hybrid vs reasoning-only models, defended benchmarks like LMSYS/Chatbot Arena, and highlighted open-model opportunities in agents, personality, and routing.

Main Topics: Tulu, RLVR, and open post-training recipes (Priority: 5/5): Lambert framed Tulu as an effort to compress complex frontier post-training into tractable, reproducible open-model recipes. He argued AI2’s work shows open models can match or beat larger labs on core evals by scaling preference data and adopting mature post-training methods. From RLHF to RLVR: verifiable rewards as the next frontier (Priority: 5/5): The conversation traced how RL from verifiable rewards emerged from math/code tasks into a general method for instruction following, precise outputs, and other checkable domains. Lambert said the naming and framing helped the idea spread, especially once major industry figures adopted it. Agentic reasoning, search, and tool use (Priority: 5/5): A major thread was how models like O3 and Deep Research combine reasoning with search and tools. Lambert argued that real progress comes from training models to use tools well, handle uncertainty, and perform multi-step retrieval/editing/search tasks rather than relying purely on end-to-end outcome rewards. Planning, abstraction, and calibration as reasoning skills (Priority: 4/5): Lambert proposed a taxonomy of reasoning skills: skills, strategy, abstraction, and calibration. He argued future agentic models need better planning, task decomposition, and compute awareness so they know when to search, backtrack, or stop rather than overthinking. Benchmarks, arenas, and evaluation strategy (Priority: 4/5): The hosts debated LMSYS/Chatbot Arena, single-turn versus multi-turn evals, and whether arenas can be gamed. Lambert defended their usefulness as community focal points and argued that better multi-turn, domain-specific, and tool-aware evaluations are needed for frontier agent behavior. Open models, personalities, routing, and product direction (Priority: 3/5): Lambert identified opportunities beyond reasoning: character/personality tuning, model spec work, and router-style systems that select among specialized models. He also emphasized open models’ appeal for private-data use cases, local deployment, and customizable behavior.

Key Arguments: Open-model post-training is increasingly about reproducing frontier practices with fewer tasks but better recipes, not inventing entirely new paradigms. RLVR is broader than math: it applies wherever outputs can be verified, including code, instruction following, and some tool-usage tasks. The main bottleneck for agents is often not the algorithm but data, curricula, and the design of tasks/verifiers. Reasoning models need skills beyond raw accuracy: planning, abstraction, and calibration determine whether they can act effectively in complex environments. Search is becoming central to reasoning because long-tail knowledge and niche information are better solved by retrieval than memorized parameters. Hybrid reasoning may be a transitional form; eventually the best models may simply learn to spend the right amount of compute dynamically. LMSYS/Chatbot Arena remains valuable despite criticism because it captures user-facing chat quality and helps the community hill-climb on a shared signal. Open models should compete on practical controllability—personality, routing, private data, and customization—not only on frontier benchmark scores.

Data Points: Tulu post-training task count: 10 to 15 tasks - Lambert said AI2’s Tulu post-training suite is relatively small compared with frontier labs’ many evals/tasks. Model scale mentioned in Tulu results: 870 and 405B - He referenced core evals for AI2 models based on Llama at those scales, saying they matched or beat Meta on key benchmarks at the time. Open preference dataset prevalence: 1 dataset over ~1 year - He noted the academic community had relied heavily on UltraFeedback since the Zephyr era, still using it as state of the art a year later. RLVR acronym adoption timing: Post-DeepSeek; after Jensen used it - He said the term took off once major industry figures started using the acronym publicly. Tool-use failure example: 80 failed attempts, success on 81st - He described RL behavior where a model repeatedly fails at tool use before eventually learning to use the tool correctly. Chatbot Arena funding: $100 million - He referenced Arena’s new funding while debating whether it changes the benchmark’s relevance. Model parallel reruns in reasoning products: 8 runs - He mentioned a theory that products like O1 Pro/DeepThink may run the model multiple times and select the best output. Open model size example: 32B - He said an Olmo 32B model is, if squinted at, roughly like original GPT-4 level while remaining fully open. Hardware price change: 4090 prices doubled in the last year - He used this as an aside on the economics of local inference and hardware ownership.

Pivotal Quotes: "Everyone just does RL on outputs." — Nathan Lambert: He said this helped crystallize the idea behind RLVR as a general method rather than a narrow math-only trick. "The model should spend the right amount of tokens on it." — Nathan Lambert: He used this to summarize the ideal of reasoning systems that dynamically regulate inference effort. "RLVR is not mature enough, nor is it as interesting of a book." — Nathan Lambert: He explained why he still plans an RLHF book rather than rebranding it around RLVR.

Implications: Open AI progress is shifting from bigger models alone to better post-training, verifiers, tools, and data. For builders, the biggest opportunities are in agent curricula, evals, and customizable open systems that can work with private data and specialized workflows.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast