The Cognitive Revolution
The Cognitive Revolution

Living Lindy: a No-BS Conversation on AI Agents with Flo Crivello

Flo Crivello, CEO of AI agent platform Lindy, provides a candid deep dive into the current state of AI agents, cutting through hype to reveal what's actually working in production versus what remains challenging. The conversation explores practical implementation details including model selecti

Featured Speakers

Nathan Labenz and Erik Torenberg HostFlo Crevello Guest

Topics Discussed

Episode Summary

Executive Summary: Flo Crevello argues that today’s practical “agents” are mostly LLM-driven workflows with deterministic scaffolding, not fully autonomous systems. He emphasizes simple, reliable tool use, heavy iteration, curated examples, and model-label abstraction over hypey multi-agent or memory architectures. Lindy’s value today is strongest in email, research, support, recruiting, and daily company intelligence, with scaffolding likely remaining important even in an AGI future.

Main Topics: What counts as an agent (Priority: 5/5): Flo favors a broad definition: software where at least part of the control flow is determined by an LLM. He stresses that agentic behavior exists on a spectrum, with most production systems still being structured workflows rather than fully open-ended autonomy. Deterministic scaffolding vs open-ended autonomy (Priority: 5/5): Lindy originally leaned too hard into fully open-ended agents, then backtracked toward deterministic steps for reliability. Flo argues that tools should usually be dumb/deterministic, while the agent makes decisions within clearly defined guardrails. Multi-agent systems and protocols (Priority: 4/5): He is excited about multi-agent futures but says they are still immature and harder to make reliable than single-agent systems with tools. He sees emerging protocols (like Google’s work and agent communication standards) as necessary infrastructure, similar to EDI in logistics. Model choice, upgrades, and evals (Priority: 5/5): Flo says model upgrades are powerful but risky because they effectively swap the ‘brain’ of AI employees. Lindy uses model labels (fastest/balanced/smartest/default) and has suffered at least one bad rollout, reinforcing the need for stronger evals and rollback discipline. Prompting, examples, and human-in-the-loop learning (Priority: 5/5): He считает few-shot examples more important than instructions and recommends starting with human-in-the-loop workflows to collect gold-standard corrections. Lindy uses user feedback and confirmations to improve behavior incrementally. Memory and RAG are becoming simpler (Priority: 4/5): Flo is skeptical of elaborate memory frameworks and thinks the bitter lesson will favor simpler approaches: inject more context directly, use straightforward scoring, and reserve RAG for cases where it truly helps. He believes many future memory systems will be lightweight and easily debuggable. AGI future and the role of scaffolding (Priority: 4/5): Even if models become far more capable, Flo expects scaffolding to remain valuable for reliability, speed, observability, and guardrails. He sees potential for AI employees with voices/faces, but believes the interface and orchestration layer will still matter.

Key Arguments: Most real-world agents today are intelligent workflows: LLMs make decisions inside human-designed control flow, rather than independently running end-to-end. Tools should generally be non-agentic; nesting agents inside tools makes systems harder to reason about and debug, and tends to perform worse empirically. Multi-agent systems are promising but immature; single-agent systems with tools and some deterministic scaffolding are currently the most reliable production pattern. Model upgrades can materially change behavior, so defaults should be abstracted and managed carefully; evals matter, but current eval suites are still weak for broad deployment decisions. Few-shot examples and human-in-the-loop correction are more effective than lengthy instructions for teaching behavior and building performance. Fine-tuning is useful only in narrow, high-volume, high-value, and sufficiently well-bounded tasks; it is not the default answer for most agent use cases. Memory systems should stay simple; as context windows and model utilization improve, many separate memory architectures may become unnecessary or too hard to debug. In an AGI/ASI future, scaffolding may not disappear; it may shift from helping models do more to helping humans understand, constrain, and safely use more capable systems.

Data Points: Flo’s appearances on the podcast: 6th appearance - The host notes this is Flo Crevello’s sixth time on The Cognitive Revolution. Agent task-length doubling time (Metaculus/METR-style curve referenced): ~7 months historically; possibly ~4 months more recently - Discussed as a controversial benchmark for how quickly agents can handle longer tasks. Meeting-video generation length in Google Gemini VO3 mention: 8-second videos - Opening sponsor read describing Gemini app capabilities. Candidate outreach workflow cost: $12 for 30 engineers / about 40 cents per engineer - Flo gives a recruiting/prospecting example using APIs to find engineers. Outreach cost per lead in the example: About $3 for follow-up outreach - Same recruiting workflow after lead generation. Daily company context ingestion: Hundreds of thousands of tokens every day - Flo describes a Lindy that ingests customer calls, support tickets, and chatbot interactions, then produces a daily digest. Model labels used in Lindy: default / fast test (Gemini 2.5 Flash) / most balanced (Claude 3.7 Sonnet) / smartest (o3) - Flo explains Lindy’s model abstraction layer and how defaults are managed.

Pivotal Quotes: "software which at least part of the control flow is defined by an LLM" — Flo Crevello: His preferred definition of an agent, credited to LangChain’s Harrison Chase. "You really should draw a pretty sharp line between your agents and your tools, and your tools should not be agentic, basically." — Flo Crevello: On why Lindy moved away from agentic tools toward deterministic tools plus an agent. "I think the scaffolding is always going to buy you something, it's always going to buy you some extra reliability, some extra speed" — Flo Crevello: His view that scaffolding remains valuable even as models improve, including in an AGI future.

Implications: For builders, the winning pattern is likely structured workflows plus good models, not magical autonomy. For companies, the biggest near-term gains are in high-volume communication, research, and operational triage. Long term, scaffolding may remain essential as both a reliability layer and a safety/visibility layer.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution