Episode Summary
Executive Summary: The conversation argues that AI agents are still early but approaching a practical tipping point. Div Gerg says open-ended agents need better data, search, verification, and domain-specific fine-tuning, while benchmarks can overstate real-world readiness. Multion is pursuing a middle path: agentic systems that are more open than scripted workflows but more constrained than fully general assistants, with 2025 framed as the likely breakout year for useful computer-use agents.
Main Topics: State of AI agents after GPT-4 (Priority: 5/5): Div describes the last 18 months as the first wave of a broader agent explosion: exciting but still glitchy, with reliability improving and real applications only now emerging. Open-ended agents vs. on-rails workflows (Priority: 5/5): The discussion contrasts fully autonomous agents with narrow, human-designed workflows. Div argues Multion is taking a middle path because consumer behavior and real tasks vary too much for rigid scripts, but fully open-ended agents are still too unreliable. Data collection and feedback for agent training (Priority: 5/5): A major theme is how to gather high-quality human-computer interaction data. Div says crowdsourced teaching data is noisy, while curated annotation and better feedback signals are more promising for improving agents. Model strategy, fine-tuning, and inference-time compute (Priority: 4/5): Div says base models are converging in capability, so performance increasingly comes from data, post-training, and inference-time reasoning/search. LoRA remains practical for smaller datasets, while larger-scale training would warrant fuller fine-tuning. Agent Q, search, and benchmark gains (Priority: 5/5): The Agent Q paper is presented as evidence that search plus learning can dramatically improve task performance. Div emphasizes Monte Carlo tree search, DPO, and trajectory-level training as the key ingredients behind large gains on narrow web tasks. Benchmarks vs. real-world performance (Priority: 4/5): The speakers note a recurring mismatch: agents can hit near-human or superhuman benchmark scores in narrow domains, yet still fail in messy, general web use. Div attributes this to the gap between mock environments and the open internet. Authentication, platform dynamics, and the 2025 outlook (Priority: 4/5): The conversation ends on agent identity/authentication, cooperation vs. resistance from websites, and Multionβs positioning against hyperscalers. Div expects a wave of practical vertical apps and assistant-like experiences in 2025.
Key Arguments: AI agents are still in an early, internet-like phase: current systems are glitchy, but reliability and adoption should rise as data and scaffolding improve. The best near-term approach is neither fully open-ended autonomy nor rigid scripts, but a middle ground with constrained tasks, human preferences, and adaptive learning. Crowdsourced correction data can help, but it is noisy; higher-quality annotation and more deliberate data collection are needed to make agents robust. There is a missing market for human-computer-use data, especially personalized/private data, but privacy, trust, and cost determine what can realistically be collected. Base model differences are shrinking; the main differentiators are task-specific data, post-training, inference-time compute, and search. Monte Carlo tree search plus learning can unlock large performance jumps because exploration discovers better trajectories that can then be reinforced. Benchmark wins should not be confused with general readiness because benchmarks are narrow and can be saturated with domain-specific tuning. Websites and services will eventually adapt to agents, but short-term incentives may be mixed, with some platforms welcoming agents and others trying to block them. The biggest constraint on adoption is not only model quality but product design: reliability, reversibility of actions, authentication, and user trust matter a lot. Multion sees its advantage in product-focused research: building specialized agentic experiences faster than frontier labs can focus on narrow verticals.
Data Points: People surveyed: 50,000+ - Div says Multion has surveyed and observed behavior from more than fifty thousand users, revealing strong variation in consumer behavior. Performance on OpenTable task with base model: under 20% - In Agent Q experiments, a Llama 3 70B instruct base model started below 20% success on a real web task. GPT-4.0 performance on OpenTable task: over 60% - The same benchmark showed GPT-4.0 doing substantially better than the base Llama model before additional tuning. Improved benchmark result after Agent Q techniques: about 95% - Div says the combination of search, feedback, and tuning pushed performance to roughly 95% in a narrow environment. Improvement factor: roughly 4.5x - Div characterizes the jump from around 20% to about 95% as a roughly 4.5x improvement. Training duration: one day / a couple of days - Div emphasizes that the gains in Agent Q came quickly, with a very short iteration cycle rather than a massive training run. Training data collected through teaching: some millions of trajectories - He says the team collected millions of trajectories via human teaching/correction methods. Model size mentioned: Llama 3 70B - The discussion repeatedly references Llama 3 70B as a key baseline model in the experiments. Preferred compromise model size: Llama 17B - Div says a 17B-sized Llama model has been a good compromise between speed and reasoning capability. Time horizon for broad agent adoption: 2025 - Both the host and Div suggest 2025 could be the year agents start becoming meaningfully useful in mainstream settings.
Pivotal Quotes: "I would call it the explosion of applications, which I don't think has happened so far." β Div Gerg: He is describing his view that the agent market is just before a broad application wave rather than in the middle of it. "The thing that's missing the most right now is just having a really great qualitative report that you can use from the environment." β Div Gerg: He explains why computer-use agents still struggle: real-world environments lack clean reward signals. "I think we're starting to see this instead of throwing compute during training, you throw computer during inference." β Div Gerg: He argues that reasoning improvements are increasingly coming from inference-time search and compute rather than only larger training runs.
Implications: Agent reliability is improving, but winners will likely be companies that combine strong data, search, and product design around specific tasks. Expect more vertical, semi-autonomous assistants in 2025, plus growing tensions over authentication and website access.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co