The Cognitive Revolution
The Cognitive Revolution

The Customer Service Revolution: Building Fin, with Eoghan McCabe & Fergal Reid of Intercom

Today Eoghan McCabe and Fergal Reid of Intercom join The Cognitive Revolution to discuss building their AI customer service agent Fin, exploring how they achieved a 65% resolution rate through rigorous optimization and custom model training rather than relying on base model improvements, while pione

Featured Speakers

Nathan Labenz and Erik Torenberg HostFergal Reed GuestOwen McCabe Guest

Topics Discussed

Episode Summary

Executive Summary: Owen McCabe and Fergal Reed of Intercom discuss the development and scaling of Finn, their AI customer service agent. They reveal that intelligence is no longer the limiting factor for customer service automation—GPT-4 was already sufficient—with over 30 points of resolution rate improvement coming from context engineering, retrieval optimization, and workflow design. Key insights include their 99-cent-per-resolution pricing model, the importance of production A/B testing over offline evals, the vision of expanding from service agents to full customer lifecycle agents, and the nuanced impact on employment (slowed hiring but not layoffs).

Main Topics: AI Strategy and Keeping up with Developments (Priority: 4/5): How Intercom stays current with fast-moving AI, including distributed responsibility across a 50+ person AI group, rapid prototyping to test new models, and reliance on trusted internal experts. The Marginal Value of Intelligence vs. Engineering (Priority: 5/5): Fergal's assessment that GPT-4-level intelligence is sufficient for most customer service work and that most gains come from retrieval, re-ranking, prompting, and workflow optimization rather than model improvements. Real-World Impact on Employment and Service (Priority: 5/5): How Finn primarily addresses underwater support teams by increasing service capacity rather than causing layoffs, with most customers pausing hiring rather than reducing headcount. Evaluation Methodology and the Role of Scale (Priority: 4/5): Intercom's approach to testing: skeptical of offline evals, relying on large-scale A/B tests in production to detect tenth-of-a-percentage-point changes, with scale as a competitive advantage. Outcome-Based Pricing (99 Cents per Resolution) (Priority: 4/5): The rationale behind pioneering per-resolution pricing: alignment of incentives with customers, initial unprofitability turned to software-level margins, simplicity, and value-based pricing principles. Future of Customer Agents and Architecture Evolution (Priority: 3/5): Expansion from service agents to customer agents covering the full lifecycle (sales, onboarding), with a shift toward more agentic architectures while maintaining building-block approaches for reliability. Internal AI Adoption and Developer Productivity (Priority: 3/5): Intercom's internal push for 2x engineering productivity via AI tools like Claude Code, balancing optimism with skepticism, and the need for structural incentives to overcome resistance to change.

Key Arguments: Intelligence is no longer the bottleneck for customer service AI—GPT-4 was already sufficient; most gains come from context engineering, retrieval optimization, and workflow design. Most customer service teams are underwater, so AI agents like Finn initially increase service capacity rather than causing layoffs, but hiring plans are affected. Offline evals are insufficient due to the messiness of real human interaction; large-scale A/B testing in production is essential for detecting small but meaningful changes. Outcome-based pricing (per resolution) creates strong alignment between vendor and customer and has become profitable despite initially losing money on each resolution. The market will consolidate around unified 'customer agents' rather than multiple specialized agents competing for the same customer touchpoints. Internal AI adoption requires structural pushes like the 2x productivity goal to overcome natural resistance to change, but must be applied pragmatically. The incumbent advantage is real but limited—new entrants with momentum in AI categories can be difficult to unseat because AI requires a fundamentally different culture and skill set.

Data Points: Resolution Rate Improvement Since Launch: ~30 percentage points - From approximately 35% to 65% over two years, with ~1% per month improvement. Contribution of Model Improvements to Resolution Rate: A few percentage points - Of the 30+ point improvement, only a small portion came from better underlying LLMs. Initial Cost Per Resolution: $1.21 - Finn's cost per resolution at launch was $1.21, but they priced at $0.99, initially losing money. Human Cost Per Resolution (Intercom): $26 - All-in cost per human-led resolution including salaries, benefits, and overhead. Intercom's Own Resolution Rate: High 60s - Intercom's own Finn deployment resolves in the high 60% range, with top customers reaching high 80s or 90s. Internal Engineering Productivity Goal: 2x - CTO set goal to double engineering output through AI tooling adoption. Finn's Model Count: 10-15 prompts - Core Finn architecture uses 10-15 distinct prompts/prompt chains. Intercom's Total Customer Base: 400,000+ - Intercom has over 400,000 businesses using its platform, providing scale for testing. Months of Sustained 1% Improvement: 30+ months - Finn has delivered ~1% resolution rate improvement per month for 30 consecutive months.

Pivotal Quotes: "Intelligence is no longer the limiting factor for customer service automation. GPT-4 was already intelligent enough for the vast majority of customer service work." — Interviewer (paraphrasing Fergal): Summary of Fergal's assessment that model capability has been sufficient for most customer service tasks, with gains coming from other engineering work. "Almost all of the improvement is in, like, you know, what you might call the rag layer or the AI layer outside the core models. Core models have definitely gotten better, but that has only contributed a couple of percentage points." — Fergal Reed: Explaining that the vast majority of Finn's resolution rate improvement comes from retrieval, re-ranking, prompting, and workflow design rather than model improvements. "The new economics and the accessibility and ease of deployment Means that before it replaces a bunch of humans, it actually Increases the supply for the things the humans were doing that one could never afford to deliver in the past." — Owen McCabe: Describing how AI customer service agents first expand service capacity rather than immediately causing layoffs, as teams address previously unmet demand.

Implications: AI product builders should prioritize context engineering, retrieval optimization, and systematic testing over chasing the latest model releases. Outcome-based pricing can align incentives and become profitable. The near-term employment impact is slowed hiring rather than mass layoffs, as AI primarily addresses existing capacity gaps.

🔓 Sign Up for Unlimited Episode Search

About The Cognitive Revolution

A biweekly podcast where hosts Nathan Labenz and Erik Torenberg interview the builders on the edge of AI and explore the dramatic shift it will unlock in the coming years. The Cognitive Revolution is part of the Turpentine podcast network. To learn more: turpentine.co

View all episodes from The Cognitive Revolution