The Cognitive Revolution
The Cognitive Revolution

AI Automation: Making AI Work for You

Nathan presents a comprehensive AI automation framework developed over three years, applicable to process automation and generative AI integration. This episode of The Cognitive Revolution, offers critical insights for businesses looking to leverage AI effectively. Learn about choosing AI tasks, und

Featured Speakers

Nathan Labenz and Erik Torenberg HostNathan Levez Guest

Topics Discussed

Episode Summary

Executive Summary: Nathan Levez presents a practical framework for AI automation: choose tasks where intelligence is genuinely required, deeply document how humans do the work, then optimize with prompting, retrieval, and fine-tuning. He emphasizes starting with task-sized, repetitive, low-risk work, using gold-standard examples, and comparing AI to real human performance rather than perfection. The talk argues AI can already deliver major ROI in routine business workflows.

Main Topics: What AI automation is vs. chat and agents (Priority: 5/5): He distinguishes interactive chatbot use (human stays in control) from true automation (delegating repeatable work to AI) and from agents (real-time delegated projects that are still unreliable today). Choosing the right work to automate (Priority: 5/5): The best targets are tasks where intelligence is required, the work is repetitive and task-sized, feedback is fast, risk is manageable, and the task is unpleasant or expensive for humans. Understanding and documenting the workflow (Priority: 5/5): Successful automation requires extracting implicit human know-how: breaking work down step by step, identifying decision points, and capturing the reasoning behind outputs. Optimization stack: prompts, retrieval, fine-tuning (Priority: 5/5): He frames performance improvement as optimizing information (prompting and RAG) and behavior (fine-tuning), often iterating through 10, 100, and 1,000-example stages. Human vs AI capability landscape (Priority: 4/5): AIs are superhuman at breadth and translation, strong in many routine tasks, but weaker than humans in depth, insight, and common-sense spatial reasoning. Business and implementation trade-offs (Priority: 4/5): He stresses comparing against actual human performance, not perfection; maximizing performance first; and using no-code tools when possible to reduce development and maintenance cost. Practical examples and ROI (Priority: 4/5): He cites Waymark, Athena, Google medical diagnosis, and Klarna to show that well-designed AI systems can produce dramatic savings, scale, and better user experiences.

Key Arguments: AI should be used where intelligence is required; if deterministic code can solve the problem, code is usually faster, cheaper, and more reliable. The most common mistake is choosing tasks that are too broad or job-sized; automation works best on well-scoped, repetitive tasks. Explicit context is essential because AIs do not know the specifics of your business, team, or process unless you provide them. Capturing gold-standard examples plus the reasoning behind them is the key bottleneck in most automation projects. Prompting can often get you far, but fine-tuning becomes valuable when you need durable behavior on a specific task. RAG is necessary when up-to-date or large amounts of external information must be injected at runtime. AI performance should be judged against human performance in practice, not an idealized perfect standard. The biggest advantages of AI are speed, cost, availability, and scalability, which can unlock qualitatively new business models. The best early wins come from low-risk, high-frequency tasks with clear standards and fast feedback loops. No-code automation is often the right default for internal business workflows because it is easier to build, maintain, and modify.

Data Points: Event audience size: more than 4,000 - Nathan says he spoke at the Adapta Summit to over 4,000 Brazilian business owners. Manaus population: 2 million people - He describes Manaus as a big city in the Amazon region with about 2 million residents. MNIST state-of-the-art accuracy: 99.7% - He cites this as the approximate human-level best performance on handwritten digit recognition. Traditional code accuracy on digit classification: 14% - Claude’s attempted rule-based solution to MNIST-like digit recognition. Perplexity-reported engineered solution accuracy: 80% - He cites a clever non-ML approach to digit recognition found via Perplexity. Google medical diagnosis chatbot result: outperformed human doctors - A Google study where a purpose-built model beat human doctors on diagnosis as judged by doctors. Klarna customer conversations handled: 2.3 million - Klarna’s AI assistant handled this many conversations in its first month live. Share of customer service conversations: two-thirds - Klarna said the assistant handled about two-thirds of all customer service chats. Full-time workers equivalent: 700 FTE - Klarna estimated the assistant did the work of 700 full-time people. Estimated Klarna profit uplift: $40 million in 2024 - Klarna’s estimate for increased profitability from the AI assistant. AI fine-tuning price for GPT-4.0 Mini: $3 per 1 million tokens - Nathan cites OpenAI’s fine-tuning cost for GPT-4.0 Mini. Inference input cost for GPT-4.0 Mini: $0.30 per 1 million tokens - He cites the input pricing for the fine-tuned model. Inference output cost for GPT-4.0 Mini: $1.20 per 1 million tokens - He cites the output pricing for the fine-tuned model. Examples needed for initial fine-tuning: 10 gold-standard examples - His recommended starting point for a task-specific dataset. Scaling rule of thumb: 10x more data per major performance step - He says moving from ~90% to ~high-90s often requires about 10x more examples each stage. Reading-automation conversion cost at Athena: under $1 - AI pipeline converts a one-hour client call into a written document for executive assistants. Prior human process cost at Athena: 4 hours and about $100 - The old manual conversion process for the same onboarding document. Content production increase at Waymark: 10x more content - Waymark’s AI-assisted process lets users create about ten times more content. Career advising offer: free one-on-one sessions - Sponsor segment for 80,000 Hours to listeners.

Pivotal Quotes: "Use artificial intelligence where intelligence is genuinely required." — Nathan Levez: Core thesis of the talk; AI should not replace deterministic code when code is sufficient. "The goal of automation is to get to the point where you can actually delegate a certain piece of work on a consistent basis to an AI system and have enough confidence that it will do the job well that you no longer have to check every single one of its outputs." — Nathan Levez: Defines the difference between chat-based use and true automation. "I cannot emphasize enough that if you can't get 10 really high-quality examples with reasoning, that you are not going to be able to automate this work successfully." — Nathan Levez: He identifies gold-standard examples and their reasoning as the critical bottleneck.

Implications: Listeners should focus on narrow, high-frequency, low-risk workflows and invest early in data collection and task understanding. The industry takeaway: AI value is already real, but success depends more on disciplined workflow design than on flashy model demos.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution