Episode Summary
Executive Summary: The podcast discusses the current state and future of AI agents, highlighting the transition from initial hype to a more grounded reality. The founder of Multion explains that while GPT-4 excels at chat, it struggles with logical deduction and long-context tasks, requiring sophisticated scaffolding for agentic applications. Key challenges include context management, online learning, and balancing performance with speed and cost. The conversation covers the shift from single-task agents to compound AI systems with parallel sub-agents, the importance of data quality and human oversight, and the need for responsible AI deployment with guardrails against malicious use.
Main Topics: Current State of AI Agents (Priority: 5/5): Despite initial hype, AI agents are still limited in capabilities. They can perform short, single-website tasks well but struggle with complex, multi-step operations, logical reasoning, and maintaining context over long interactions. Architectural Innovations and Scaffolding (Priority: 4/5): Multion uses a combination of fine-tuned open-source models, vision encoding, and proprietary memory management (e.g., storing and retrieving context like a CPU) to optimize performance, reduce token usage, and enable efficient large-context operations. Cost and Efficiency Trade-offs (Priority: 3/5): Running agents on GPT-4 is prohibitively expensive for consumer products. Multion focuses on internal models and caching to bring cost per step down to 2-10 cents, emphasizing speed (10x human) as a key value proposition. Future Directions: Compound AI and Multi-Agent Systems (Priority: 4/5): The roadmap includes moving from single agents to orchestrated systems with parallel sub-agents (inspired by operating systems), scheduled tasks, and mobile integration. The goal is to create an invisible 'AI scheduler' that manages complex workflows. Responsible AI and Guardrails (Priority: 5/5): The discussion covers the importance of building guardrails like prompt injection detectors, verification steps before executing actions, and using moderation APIs to prevent malicious use. The industry must self-regulate to avoid external regulation and maintain trust. Economic and Societal Impact (Priority: 4/5): AI agents will likely first complement human workers by automating 'digital chores,' then gradually replace certain jobs (like typewriter operators), creating new roles in agent management and programming. This could democratize access to assistance for those who cannot afford human help. Roadmap and Adoption Strategy (Priority: 3/5): Multion plans to start with single short tasks, progress to long tasks, then compound tasks, and finally parallel operations. Adoption will come through an API, mobile interface, and partnerships, targeting early adopters like executive assistants.
Key Arguments: Current AI models like GPT-4 are good at generating plausible-looking content but lack deep logical reasoning and can easily be confused by long context, limiting their effectiveness for complex agents. Multion's approach combines external memory, efficient prompting (~5000 tokens per step), and vision encoding to manage context effectively, achieving a 10x speed advantage over human performance. Running agents on GPT-4 is not cost-effective for consumer products; using fine-tuned open-source models (with a cost of 2-10 cents per step) is necessary for scalable deployment. The future of AI agents lies in compound systems where a 'scheduler' orchestrates multiple sub-agents in parallel, inspired by operating system kernels, enabling complex multi-step tasks. To prevent malicious use, agents need verification steps before executing actions, prompt injection detectors, and moderation classifiers. The industry should adopt minimal standards like OpenAI's moderation API. AI agents will first complement human workers by automating tedious digital tasks, but will eventually transition jobs toward agent management and programming, similar to how computers replaced typewriters. High-quality private data and human-in-the-loop training pipelines are essential for building models that can handle agentic tasks, as open-source data alone is insufficient.
Data Points: Average prompt token size per inference step: Not more than 5,000 tokens - Multion optimizes context by mixing language metadata and visual inputs, keeping prompts small to maintain model focus and reduce cost. Cost per step: 2-10 cents - Using efficient fine-tuned models, each agent action costs between 2 and 10 cents. A 100-step task costs $2-10, making it viable for consumer use. Number of steps for a task: 10-100 steps - Simple tasks (e.g., putting items in a cart) take ~20 steps; complex research or compound tasks can take up to 100+ steps. GPT-4V token equivalence for low-res input: 85 tokens - Compared to raw HTML (tens of thousands of tokens), using vision encoding is much more token-efficient and provides higher semantic density. Average cost reduction target for consumer product: GPT-4 too expensive; targeting 2-3 cents per step - GPT-4 pricing makes it impossible to support millions of users, so Multion focuses on custom models and caching for cost reduction.
Pivotal Quotes: "Just because the technology is not there, we are using humans as a substitute with us... when computers replace typewriters, it actually ended up creating more jobs, but it definitely changed the nature of the jobs." — Div Gerg: Discussing the future economic impact of AI agents: they will eliminate 'shitty jobs' and create new roles in agent management, similar to the transition from typewriters to computers. "Even if you think about language models, we just call them language models, but there's nothing inherent about them which is like, okay, this can only work on language... there's nothing about transformers where, like a transformer, people think it's a language model, but you can use it for pretty much anything." — Div Gerg: Explaining the vision for building 'action transformers' and moving beyond pure language models to handle trajectory-level optimization and multimodal tasks in agent systems. "I think it's probably in their quadratically plans now. Like maybe in the next six months, we want to have some sort of strategy around agents. I don't think people have tried agents themselves." — Div Gerg: On the current state of the market: outside the Bay Area, awareness of AI agents is still low, but many businesses are beginning to plan for agent-based automation.
Implications: AI agents are transitioning from hype to practical utility, with early adopters benefiting from 10x speed on tasks like research and e-commerce. However, reliable multi-step agents remain a research challenge. The industry must invest in guardrails and cost-efficient models to avoid a backlash from malicious use and to make agents accessible to consumers. Expect a shift toward compound AI systems within 12-18 months.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co