Episode Summary
Executive Summary: At Latent Space Live, Graham Newbig surveyed the state of LLM agents in 2024, arguing that coding/web agents are becoming practical but still need better planning, search, evaluation, and human interaction. He highlighted OpenHands’ design choices, benchmark progress, and a strong belief that 2025 will be the year agents become cheaper, more capable, and more widely deployed.
Main Topics: State of LLM agents in 2024 (Priority: 5/5): The talk frames 2024 as a breakthrough year for agents, especially coding agents, with major product launches and benchmark gains across software development, IDE assistance, support, and search. Agent-computer and human-agent interfaces (Priority: 5/5): Newbig discusses how agents should interact with computers and users, favoring a small set of powerful tools, text-first workflows, and interfaces that expose enough detail for auditing without overwhelming users. Model choice and agent reliability (Priority: 5/5): He argues that agentic systems need strong instruction following, tool use, environment understanding, and error recovery, and says Claude currently performs best in his framework. Planning, workflows, and memory (Priority: 4/5): The talk compares curated plans, on-the-fly planning, single-agent vs multi-agent systems, and workflow memory approaches that reuse successful task patterns to improve future performance. Exploration, search, and evaluation (Priority: 4/5): He covers repo/web exploration, tree search, and benchmark design, emphasizing that agents need better ways to inspect environments and that current benchmarks are becoming too easy. Open source and accessibility (Priority: 4/5): Newbig closes with a call for open source, cheaper models, and broader access so agent capabilities do not concentrate power but instead expand what more people can do.
Key Arguments: Coding agents are already useful in daily work for data analysis, script generation, and iterative software improvement. A small toolset plus arbitrary Python execution is often more effective than exposing many granular tools directly to the model. The best agent systems should provide enough transparency for users to audit actions while still keeping interaction simple. Claude currently appears strongest for agentic coding tasks because it handles instruction following and error recovery better than alternatives. Most agent failures come from insufficient information gathering before acting, not random mistakes. Single-agent systems with lightweight planning can outperform more rigid multi-agent architectures because they are easier to adapt when tasks deviate from the plan. Workflow memory and self-improvement can boost performance by reusing successful task patterns from prior runs. Benchmarks like SWE-bench and WebArena are useful but will likely saturate, requiring harder and more realistic successor benchmarks. Agents will need redesigned websites, APIs, and authentication systems to work reliably at scale. Open source models and tools are important to make agent capabilities affordable and broadly accessible.
Data Points: Latent Space Live attendance: 200 in person - Audience present throughout the day at NeurIPS 2024 in Vancouver Latent Space Live online viewership: 2,200 live viewers - People watching the mini-conference online Survey respondents: over 900 - Listeners who told the hosts what topics they wanted covered OpenHands SWE-bench full leaderboard: 29% - End-of-year position on the hardest SWE-bench full leaderboard OpenHands SWE-bench verified: 53% - Smaller verified benchmark score mentioned in the intro Amazon Q Devlo SWE-bench verified: above 53% - Named as ahead of OpenHands on the verified benchmark OpenAI O3 self-reported SWE-bench verified: 71.7% - Referenced as a higher self-reported result WebArena improvement from workflow memory: 22.5% increase - Reported gain after 40 examples using agent workflow memory Claude evaluation ranking: best in framework - Compared against GPT-4o, o1-mini, Llama 3.1-405B, and DeepSeek 2.5 Open source model ranking: Llama 3.1-405B was best open model - Old evaluation cited as the strongest open-source model at the time Human-like success estimate on own repos: 30% to 40% - Speaker’s estimate of tasks solved without human intervention on his own repositories Tasks solved with feedback: 80% to 90% - Estimate for tasks agents can solve when the human provides feedback Agent workflow memory examples: 40 examples - Number of examples after which the WebArena improvement was observed
Pivotal Quotes: "2025 is going to be the year of agents" — Charlie / intro narration: Opening framing of the talk’s broader industry outlook "If you have a really, really good instruction following agent, it will follow the instructions as long as things are working according to your plan." — Graham Newbig: Explaining why lightweight single-agent systems can work well "I think the biggest thing that it fails at is insufficient information gathering before trying to solve the task." — Graham Newbig: Describing the most common failure mode of current coding agents
Implications: Agents are moving from demos to practical tools, but reliability, UX, and infrastructure still need work. Expect cheaper, stronger models, more API-first systems, and growing pressure to make agent tech open, auditable, and accessible.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast