Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

The Age of Async Agents — Cognition's Walden Yan & OpenInspect's Cole Murray

The new AIEWF website is live! CFPs close in 2 days and we will run our first New Engineer Orientation this weekend, get your tickets booked ASAP as they -will- sell out. Take the AI Engineering Survey and get >$2k in credits and free AIE WF tickets! One of the central tensions in the agents indu

Featured Speakers

Latent.Space HostWalden Yen GuestCole Murray Guest

Topics Discussed

Episode Summary

Executive Summary: The discussion centered on the rapid maturation of background/cloud coding agents and the infrastructure needed to make them practical, secure, and useful at scale. Walden Yen and Cole Murray compared Devin/OpenInspect architecture, debated in-box vs out-of-box agent design, repo setup, testing, integrations, memory, and multi-agent orchestration, concluding that real value comes from tightly integrating agents into company workflows and keeping humans in the loop for code quality and boundaries.

Main Topics: Rise of background agents and autonomy (Priority: 5/5): The speakers argue that model capability jumps in late 2025 made it realistic to go from specification to completed pull request with minimal hand-holding, pushing teams toward cloud/background agents. Agent architecture: in-box vs out-of-box (Priority: 5/5): A major theme was whether the agent’s brain should run inside the sandbox or separately in a control plane. They favored out-of-box for security and reuse, despite higher complexity. Repo setup, environment management, and infrastructure (Priority: 5/5): They emphasized that the hardest problem is not code generation but getting the environment, secrets, dependencies, and VMs/sandboxes configured so agents can actually run and test code. Testing, screenshots, and GitHub workflow (Priority: 4/5): Testing was framed as broader than computer use: agents must orchestrate systems, trigger behaviors, inspect screenshots/videos, and interact naturally on GitHub without looping or spamming. Integrations, MCP, and company-wide adoption (Priority: 4/5): Both speakers stressed that background agents only become valuable when integrated into Slack, logs, databases, knowledge bases, and support workflows; MCP alone is often insufficient. Memory and knowledge management (Priority: 4/5): Memory was described as unresolved: useful systems need careful generation, retrieval, pruning, and possibly file-system-like skills or Markdown-based persistent context. Multi-agent systems and future workflow patterns (Priority: 3/5): They were cautious about swarms, leaning toward manager/sub-agent patterns and isolated sessions rather than chaotic peer-to-peer agent networks, though they saw growing potential for richer collaboration.

Key Arguments: Background agents became much more practical once frontier models could autonomously drive from spec to PR with little friction. Running the agent out of the box is preferred for security because secrets stay off the sandbox machine, though this increases state-management complexity. The real hard problem is not just agent intelligence; it is repo setup, dev-environment fidelity, secrets handling, and VM/sandbox lifecycle management. Testing is a broader orchestration problem than computer use: agents must run apps, trigger flows, validate outputs, and sometimes coordinate multiple frontier models. Open source and customization matter because companies want to fork and adapt agent infrastructure rather than pay for a generic per-seat product. MCP is useful but often too one-way and too limited for deep company integration; first-party, bidirectional integrations often work better. Memory remains unsolved because both retrieval and generation are tricky; low-noise, high-signal memories and pruning are essential. Humans still matter for architecture, code review, and boundary-setting; without guardrails, AI coding can regress codebases toward the worst engineering patterns. Multi-agent systems are likely to emerge, but today the most practical approach is still isolated sub-sessions with clear task segregation rather than free-form agent swarms.

Data Points: Devin commit share on repos: 16% in January to 80% in March - Walden cited this as evidence of the sharp increase in autonomous agent contribution. Merged PR usage growth: 7x - Walden referenced internal growth in merged PR activity over a short period as agents became more capable. Engineer headcount growth: about 10% - Mentioned alongside the 7x PR growth to show usage rose far faster than headcount. Model capability shift timing: December 2025 - Described as the period when Opus 4.5 and GPT 5.2 crossed a threshold enabling more autonomous PR generation. Historical model milestone: Sonnet 3.7 - Cognition reportedly rewrote Devin in one night after this jump in capability. Autonomous vibe-coding benchmark: about two weeks - They estimated December-era systems could be run with auto-merge/no code review for roughly two weeks before the codebase degraded. Typical enterprise agent spend: $1,000 to $5,000 per engineer - Estimate given for responsible and value-producing usage of AI coding agents.

Pivotal Quotes: "We could pretty much go from a specification to a completed pull request, assuming the spec was good enough with very little friction." — Walden Yen: Explaining why background agents became practical after recent model improvements. "The hardest part of the perennial problem since the start of the company of how do we help people get the setup?" — Walden Yen: On repo setup and why environment management is the central challenge for agents. "Your code base regresses to your worst engineer." — Cole Murray: Describing how unchecked AI coding patterns can entrench low-quality architecture and style.

Implications: Agent adoption will hinge less on raw model intelligence and more on infra, integrations, testing, and governance. Companies that standardize boundaries, local testability, and workflow integration will get the most leverage, while teams without guardrails risk rapid codebase decay.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast