Episode Summary
Executive Summary: Andrew Lee discusses Tasklet, an AI agent platform built to turn open-ended natural-language goals into recurring automations and ad hoc assistance. He argues that models should be trusted more than rigid workflows, that long-lived “virtual employees” are emerging, and that speed—not proprietary moats—is the key competitive advantage. The episode explores Tasklet’s architecture, integrations, memory, reliability, cost, security, and the broader future of agentic AI.
Main Topics: Tasklet’s product vision and origin (Priority: 5/5): Tasklet emerged from Shortwave users wanting automations that could run without supervision. Lee expanded the idea from email-centric workflows into a general-purpose agent platform that blends chatbot-like interaction with automation outcomes. “Betting on the model” vs workflow automation (Priority: 5/5): Lee argues traditional workflow tools are increasingly inferior because models are getting smarter. He believes fully agentic systems are more robust to real-world messiness and error states than step-by-step flowcharts. Two-tier agent architecture and long-lived relationships (Priority: 5/5): Tasklet separates a main agent from sub-agent runs, allowing users to keep chatting with an ongoing agent while it also handles recurring tasks. This is positioned as the foundation of a future virtual employee model. Context, memory, and compaction (Priority: 4/5): A major technical challenge is preserving the illusion of long-term memory without overflowing context windows. Tasklet uses SQL-backed state, context compaction, system messages, and tool-based retrieval rather than relying on external memory vendors. Integrations, MCP, and direct API connections (Priority: 4/5): Tasklet supports thousands of tools via integration platforms, MCP, direct API generation, and computer use. Lee is skeptical that MCP is always necessary, arguing direct API connections often work better and are used more frequently. Model selection, evals, and reliability (Priority: 4/5): Lee says model choice is still “vibes-based,” with Anthropic preferred for long multi-turn agent work. He prioritizes shipping fast over formal evals, using internal testing and staged rollouts instead. Safety, security, and enterprise trust (Priority: 3/5): Although Tasklet is early and not yet SOC 2 certified, Lee says trust and compliance will matter more as the product moves upmarket. He sees agent safety as a new category beyond traditional SaaS security controls.
Key Arguments: Models should be the primary source of control; software should wrap models, not the other way around. Fully agentic systems can be more reliable than workflow software because they can route around unforeseen failures at runtime. The major advantage of Tasklet is not just automation but an ongoing relationship with a persistent agent that can learn from feedback. Context engineering is a core discipline: long-lived agents need careful retrieval, compaction, and state management to simulate a long chat. Direct API connections often outperform MCP because modern models can infer API usage from docs and use tools without bespoke wrappers. Anthropic models are preferred for multi-turn agent loops because small per-turn advantages compound over dozens or hundreds of turns. Evals are less important than rapid iteration in a fast-changing product and model environment; shipping speed is the primary moat. Tasklet’s future is as a virtual employee platform, especially for business operations and recurring workflows. Security and compliance matter, but current customers prioritize capability and intelligence over certifications in the early stage.
Data Points: Tasklet launch timeline: Built from end of May/early June through launch last week - Lee says the team moved quickly from the Shortwave MCP use cases into the new platform over several months. Integration count: 3,000+ business tools - Tasklet claims broad connectivity across integration platforms, direct APIs, MCP, and computer use. Context window: 200,000 tokens - Mentioned as the current context budget for the main agent. Run limit: 50-turn limit - Tasklet initially imposed this to avoid runaway costs, but the team says they now hit it frequently and need to raise it. Model preference: Anthropic models for long multi-turn interactions - Lee says Anthropic is best for iterative LLM-call/tool-call loops, especially over many turns. Price signal: GPT-5 is described as less than half the price of Sonnet 4.5 - Lee uses pricing as evidence that Anthropic is still the preferred choice for these tasks. Human-time benchmark: A couple hours of work / ~50% of task-length frontier - Referenced as the current frontier on autonomous task duration, though Lee argues the benchmark is still evolving. Cache hit rate in Shortwave: ~85% - Lee cites this as a result of improved caching strategies. Tasklet revenue pace: Adding revenue faster than Shortwave ever did - Lee says early Tasklet revenue growth is strong despite high token costs. Cost position: Strongly margin negative - Tasklet is currently expensive to run, especially because recurring jobs consume many tokens and computer use is costly. Model rollout preference: “Still running on vibes” - Lee says model selection and upgrades are driven by hand testing and customer demand, not formal evals.
Pivotal Quotes: "“You should always bet on the model.”" — Andrew Lee: His core argument for replacing workflow-first automation with agent-first automation. "“The only thing that matters at this point is speed.”" — Andrew Lee: On competitive advantage in AI startups and why Tasklet moves quickly rather than over-investing in static moats. "“I think the future is... a virtual employee.”" — Andrew Lee: Describing the long-term product vision for persistent, long-lived agents that users revisit over time.
Implications: Tasklet suggests a near-term future where users delegate recurring work to persistent agents that improve through feedback. For builders, the message is to optimize for speed, model quality, and context engineering—not rigid workflows or static moats.
From the Transcript
Yeah, traditionally, the way people do these sorts of automations is they use a workflow product, something like Zapier or N8N or like the Agent Kit product that OpenAI put out, I would call a workflow product. And I look at workflow products and say, this was the right way to approach this from a traditional software engineering standpoint, like a year or two ago when models were smart, but not that smart. And I think we've learned over the last few years that you should always bet on the model. The models are always going to get smarter. And the right thing to do is to find ways to give those models more agency over time. And I think the reason people have been shy about doing workflow automation like fully agentic all the way down is because they didn't really trust it to be reliable. They didn't feel like the models were there yet. But I think now's the time. I think we've gone through that transition from you know, you have a workflow that defines step one, step two, step three, step one.
No, I think speed is the only thing that matters at this point. I think there used to be both, and AI is like, I'll give you a good example here: with all the integration platforms, it used to be that to compete with a Zapier or an N8N or a pipe training or something, you needed to have a lot of connectors, and those connectors had to be hand-rolled. And so you had to build up over the course of years this huge investment in building those connectors. With Direct API, that's gone, right? So we, on day one, have, I'd argue with computer use. Better integration support than Zapier has, despite being a few months old, the product being a few months old, and Zapier being a very old product. So that moat to them is gone. Shortwave, I give you a similar example of like for a long time where, hey, our moat is like, we have an email client, right? Anyone who wants to come and build an AI email client, they first have to have an email client. Well, it's not going to be too long before you can just ask the AI to write an email client for you. And then that mode is gone. So I do think it's speed.
We could build the ability for you to tell the LLM, hey, during this phase, you must do these steps in order. And the LLM could actually create its own guardrails and enforce that through code if it wanted to. So I don't, I think if the model is wrapping the code, the model can then construct constraints in those codes in that code to enforce the reliability goals that you want. I haven't built it yet, but I think we can. So I think the future is not going to be forever. Here's the Quick, easy, unreliable way, and then here's the harder, more reliable way. I think it really is going to be better in basically all scenarios in the not too distant future. Yeah, maybe it could be Sonnet 4.5 new, or maybe we'll have to wait till Sonnet 4.7. But it seems like these things are coming at us pretty fast. Hey, we'll continue our interview in a moment after a word from our sponsors. The worst thing.
About The Cognitive Revolution
A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co