Episode Summary
Executive Summary: The conversation centers on how AI agents will reshape enterprise work, with Box CEO Aaron Levie arguing that agents need governed workspaces, permissions, and identity boundaries to safely access corporate content. The guests conclude that knowledge work, unlike coding, is harder because of messy data, context rot, and high-stakes output quality—making workflow redesign, evals, and agent-ready infrastructure essential.
Main Topics: Agents Need Enterprise Infrastructure (Priority: 5/5): Levie argues that agents need a governed 'box'—a workspace with access controls, storage, and oversight—to operate safely in enterprise environments. Identity, Permissions, and Security for Agents (Priority: 5/5): A major theme is that agents cannot simply be treated like human users; they create new liability, privacy, and access-control problems that existing RBAC models do not fully solve. Why Knowledge Work Is Harder Than Coding (Priority: 5/5): The speakers contrast AI coding with broader knowledge work, noting coding has more structured data, better feedback loops, and fewer access-control constraints than enterprise workflows. Context Engineering, Search, and Context Rot (Priority: 5/5): They discuss retrieval challenges, limited context windows, search quality, pruning mistakes from context, and the need for agents to know when to stop searching. Evaluations and Agent Quality Measurement (Priority: 4/5): Box’s internal evals and public partnerships like Apex Agents are presented as critical for measuring model and agent improvements across industries and workflows. Workflow Transformation and Organizational Change (Priority: 4/5): The discussion emphasizes that companies must re-engineer processes to make agents effective; agents do not just drop in and automate work without operational changes. Box as a Read-Write Workspace and Future Knowledge Base (Priority: 4/5): Beyond read/search use cases, Box sees agents creating and storing files, memory, and working docs in sandboxed enterprise workspaces, potentially including authoring workflows.
Key Arguments: Agents are not adapting to human work; enterprises are adapting their work to the agent model. Every agent needs a box: a governed workspace with permissions, storage, and oversight. Enterprise adoption will require new identity, liability, and access-control frameworks because agent accounts are not equivalent to human accounts. Coding is an unusually favorable AI use case because it is text-heavy, technically fluent, and highly instrumented; other knowledge work lacks those advantages. Search and retrieval remain unsolved at scale because agents face a 10 million documents vs. 60,000 tokens problem. Models are improving at judging when a result is wrong or incomplete, but they still struggle in messy data environments. Companies will need to become more documented, structured, and agent-ready, or competitors using agents will gain compounding productivity advantages. Evals will become mandatory across enterprises for AI-generated work products such as RFPs, sales materials, contracts, and invoice processing. Agents will likely force a new market for consulting and DevRel/FDE-style roles that help companies redesign workflows and adopt agentic systems.
Data Points: Box customer penetration among Fortune 500: 67% - Levie says Box serves 67% of the Fortune 500. Internal Box company size: ~3,000 people - Described while discussing Box’s core agent team inside the broader company. Core AI/agent team size: a few dozen people - Levie describes a close-knit epicenter working on agents, evals, search, and supporting systems. Common enterprise ramp time: 1 to 3 months - Used to argue that much workplace knowledge is tacit and not documented. Potential improved onboarding via documentation: 3-month ramp to 2-week ramp - Presented as an example of productivity gains from agent-ready documentation. Context window for accurate work: about 60,000 tokens - Referenced as a practical working range for reliable retrieval despite larger theoretical windows. Corporate corpus size example: 10 million documents / 50 million pages - Used to illustrate the gap between available knowledge and model context limits. Office-address benchmark: 10 offices - Example internal benchmark where an agent must find addresses for 10 Box offices. Model improvement example: 15-point jump - Levie mentions a roughly 15-point overall jump in Box’s internal eval between model generations. Box AI/agent workflow timing: 11 p.m. - He jokes that he often ends up in late-night Zooms discussing agents and platform features.
Pivotal Quotes: "Every agent needs a box." — Aaron Levy: Summarizing Box’s thesis that agents need a governed workspace to function safely in the enterprise. "We are changing our work to make the agents effective in that model. The agent didn't really adapt to how we work. We basically adapted to how the agent works." — Transcript/guest framing: Core argument about workflow redesign and the direction of adaptation in AI adoption. "You don't write code, you talk to an agent, and it goes and does it for you." — Aaron Levy: Describing how coding workflows have been transformed by AI agents.
Implications: Enterprises will need governed agent workspaces, stronger documentation, better search, and rigorous evals. The winners will be companies that redesign workflows early and build infrastructure for safe, auditable agent collaboration.
About Latent Space: The AI Engineer Podcast
The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space
View all episodes from Latent Space: The AI Engineer Podcast