Episode Summary
Executive Summary: Aaron Levy argues the headline that “95% of companies get no return on AI” is misleading because AI is still early, most deployments are pilots or DIY attempts, and real ROI comes from applied solutions, good context, and workflow redesign. He says agents are emerging as the core model for knowledge work, with Box building workflow-based AI tools that use enterprise content to automate tasks like extraction, review, and document generation.
Main Topics: Why the MIT AI ROI study is misleading (Priority: 5/5): Levy rejects the study’s broad conclusion, arguing it mixes early pilots, poor DIY implementations, and immature measurement with real enterprise value already being seen in applied AI use cases. DIY AI vs applied enterprise solutions (Priority: 5/5): He says companies trying to build full AI stacks internally often fail, while purpose-built external solutions work better because they reduce complexity and integrate security, permissions, and context. AI changes workflows, not just tasks (Priority: 5/5): Levy argues enterprises must re-engineer processes to get real gains from AI; AI does not simply slot into old workflows, especially for coding, sales, legal, and operations. What an AI agent is (Priority: 5/5): He defines agents as systems that do work for you by looping through models multiple times, maintaining memory, and executing multi-step tasks, with coding as the clearest current example. Box’s product strategy: Box Automate and agents (Priority: 4/5): Box is rolling out workflow-based tools that combine enterprise content with AI agents to automate tasks such as contract extraction, proposal generation, onboarding, and document review. GPT-5 and model progress (Priority: 4/5): Levy says GPT-5’s reception was initially muted because users had already seen many intermediate model jumps, but Box’s evaluations show meaningful improvements in enterprise document and reasoning tasks. Economic stakes and AI investment (Priority: 4/5): He argues AI losses and huge capital spending are rational because the prize is automating knowledge work across major industries, creating a potentially enormous economic payoff.
Key Arguments: The MIT-style ROI story is distorted because early AI adoption is dominated by pilots, experimentation, and poor-fit DIY builds rather than mature deployments. External, applied AI products outperform internal builds because most non-tech companies lack the expertise to assemble secure, scalable AI stacks. AI ROI depends on changing workflows and roles; workers increasingly become managers/reviewers of AI output rather than direct executors of all tasks. Enterprise value is strongest when AI is grounded in high-quality, trusted context from existing company data. Agents are not just chatbots; they are multi-step systems that can loop, remember, and complete work over minutes or hours. The best evidence of AI value is visible in coding, where small teams can now operate with output comparable to much larger teams. Consumer AI feels slower than enterprise AI because packaging trustworthy, affordable, on-device experiences for mass users is far harder than demoing frontier-model capability. GPT-5 is best understood as an incremental culmination of prior model gains, not a shocking leap, though Box’s enterprise evals still found clear improvements. Massive AI spending is justified if the technology truly automates knowledge work at scale across healthcare, law, finance, engineering, and other sectors.
Data Points: Organizations with zero return on AI investment: 95% - Figure cited from the MIT study discussed in the episode. Enterprise investment in generative AI: $30 billion to $40 billion - MIT study estimate of spending on generative AI. Internal build failure rate vs external partnerships: Internal builds fail at double the rate - Referenced from the study and used to support Levy’s argument against DIY AI stacks. Employees using personal AI daily: 90% - Surveyed employees reportedly use personal AI daily even when official LLM purchasing is lower. Official LLM purchases by firms: 40% - Study noted that only a minority of firms have formal LLM purchases. Box customer base: 120,000 customers - Used to explain why Box can leverage open source and AI expertise at scale. ChatGPT era duration: Nearly 3 years - Levy notes that despite this time, reliable document generation is still emerging. AI agent adoption timeline: 2025 is the first year - He says 2025 is the first year it is serious to talk about agents in this way. OpenAI cash burn through 2029: $115 billion - Used in discussion of the economics of AI infrastructure investment. Oracle deal size: $300 billion - Mentioned as a major cloud/infrastructure commitment tied to AI growth. Startup team efficiency example: 9-person startup operating like a 100-person company - Levy cites a startup using AI coding tools to illustrate productivity amplification.
Pivotal Quotes: "We are still early in the adoption curve of AI." — Aaron Levy: Used to explain why many AI projects are still pilots or experimental and why ROI data is noisy. "We will actually have to re-engineer some of our business processes to make agents effective." — Aaron Levy: Levy’s core point that AI requires workflow redesign, not just software insertion. "We will be the reviewers of the AI agents' work. We will be the editors. We will be the managers. We will be the orchestrators." — Aaron Levy: Describing how knowledge work roles will shift as agents take over more execution.
Implications: Listeners should expect AI value to come from redesigned workflows, strong context, and agent-driven automation—not generic copilots. Enterprises that adopt applied solutions early may gain major productivity advantages, while laggards risk disruption.
About Big Technology Podcast
The Big Technology Podcast takes you behind the scenes in the tech world featuring interviews with plugged-in insiders and outside agitators. Alex Kantrowitz, a Silicon Valley journalist who's interviewed the world's top tech CEOs — from Mark Zuckerberg to Larry Ellison — is the host.