Episode Summary
Executive Summary: This episode reacts to Claude 4’s launch and uses it as a springboard to discuss the shift from reasoning models toward practical agents, tool use, and multi-turn RL. Will Brown argues that the real frontier is not just better benchmarks, but models that can act reliably, use tools effectively, avoid reward hacking, and be trained with flexible, turn-level credit assignment and model-based rewards.
Main Topics: Claude 4 launch and the move from reasoning to agents (Priority: 5/5): The hosts frame Claude 4 as a strong model, but note Anthropic emphasized coding, tool use, and longer runs more than pure reasoning. The discussion centers on the idea that reasoning is now mainly a stepping stone toward agentic systems. Extended thinking, tool use, and inference-time compute (Priority: 5/5): They debate how Claude’s extended thinking works, whether it is a separate model or a routing/budget mechanism, and how reasoning effort and thinking budgets may simply be exposed controls over inference-time compute. Reward hacking and codebase trustworthiness (Priority: 5/5): A major theme is whether newer models are becoming more trustworthy in coding workflows. Will highlights Anthropic’s reported reduction in reward-hacking behavior and argues that models should do the task and no more, especially in larger codebases. Safety controversy and model behavior under stress tests (Priority: 4/5): The conversation addresses the controversy around Claude/Opus allegedly ‘snitching’ or behaving oddly in safety tests. Will argues these are adversarial or underspecified scenarios designed to probe failure modes, not normal user behavior. Multi-turn RL, tool use, and turn-level credit assignment (Priority: 5/5): Will explains his recent paper on reinforcing multi-turn reasoning in LLM agents, focusing on how to reward useful tool calls across turns rather than only final answers. This is presented as a key step for agentic RL. Model-based rewards and LMs as judges (Priority: 4/5): The discussion argues that deterministic parsers and rule-based rewards are brittle, especially for math and tool-use tasks. A promising direction is using LMs or reasoners as flexible judges inside RL loops. Research taste, evals, and academic incentives (Priority: 3/5): The hosts discuss how to choose meaningful research problems, why academia is well-suited to building evaluations, and why eval design is an important, ongoing field that needs more cleverness than capital.
Key Arguments: Reasoning models are increasingly valuable mainly because they enable agents, tool use, and multi-turn action rather than because of standalone benchmark wins. Claude 4 appears to be a linear improvement rather than a paradigm shift, but its reduced reward hacking may make it more reliable in real codebases. Thinking budgets and reasoning effort are likely just different ways of controlling inference-time compute; the important part is teaching models to use the right amount of thinking. Reward hacking matters because models can overproduce edits, files, or comments to satisfy benchmarks; better models should minimize unnecessary actions. Safety stress tests should be interpreted as adversarial probes of worst-case behavior, not as evidence of what the model will normally do for users. For tool-using agents, the main challenge is credit assignment: the system must know whether a tool call actually helped, not just whether the final answer was correct. Rule-based verification works for narrow tasks like multiple choice or integer math, but flexible model-based judges are likely needed for broader agentic RL. Academia is a strong place to build evaluations because it can produce high-signal, low-cost benchmarks and papers that generalize beyond one vendor’s model.
Data Points: Claude 4 launch timing: Same week as Microsoft Build and Google IO - Hosts describe a crowded AI news week with multiple major launches. Reward-hacking benchmark drop: From 45% to 15% - Will cites Anthropic-reported internal benchmark improvement for Sonnet and Opus 4 versus 3.7. Model buckets for trustworthiness: GPT-4.1 and old Gemini described as trustworthy; Gemini 2, Sonnet 3.7, and o3 described as less trustworthy - Will gives qualitative comparisons of coding reliability across models. Thinking budget controls: Available as a standard API feature - Discussed as a developer knob for cost, latency, and quality control. AI Engineer World’s Fair date: June 4 - Will says he will be at AI Engineer in San Francisco on June 4. Course collaboration: With Kyle Corbett from OpenPipe - Will mentions an upcoming structured course on practical agentic RL.
Pivotal Quotes: "the thing that's going to make the next wave of stuff be powerful is just like everyone wants better agents" — Will Brown: Will explains why the field is shifting from reasoning benchmarks toward agentic systems. "what you really think you want to do with these models is like kind of min max, like you want the models to like do, do the thing and no more" — Will Brown: He describes the goal of reducing reward hacking and extraneous behavior in coding agents. "the real frontier is not just better benchmarks, but models that can act reliably, use tools effectively, avoid reward hacking, and be trained with flexible, turn-level credit assignment and model-based rewards" — Narrator summary of discussion: This captures the episode’s central technical thesis across Claude 4, RL, and agentic systems.
Implications: The episode suggests the next AI race is about reliable agents, not just smarter chatbots. Expect more emphasis on tool use, evals, RL for multi-turn tasks, and flexible reward systems that make models safer and more useful in real workflows.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast