Latent Space: The AI Engineer Podcast
Latent Space: The AI Engineer Podcast

Notion’s Token Town: 5 Rebuilds, 100+ Tools, MCP vs CLIs and the Software Factory Future — Simon Last & Sarah Sachs of Notion

For all those who missed out on London, see you in Miami next week! Notion, the knowledge work decacorn, has been building AI tooling since before ChatGPT, with many hits from Q&A in 2023 and unified AI in 2024 and Meeting Notes in 2025. At the end of their last Make user conference, Ryan Nystro

Featured Speakers

Latent.Space Host

Topics Discussed

Episode Summary

Executive Summary: Notion leaders Simon and Sarah discuss how custom agents, evals, and a “software factory” are reshaping Notion into an AI-native system of record. They argue the company wins by shipping practical agent workflows, not just demos, while balancing MCP, CLIs, pricing, and rapidly changing model capabilities.

Main Topics: Custom agents as the core product (Priority: 10/5): Notion’s new agents automate real work like inbox triage, bug routing, and leases. From model limits to product strategy (Priority: 9/5): The team learned to stop forcing model fits and instead redesign workflows around capabilities. Evals and model behavior as infrastructure (Priority: 10/5): Notion treats evals like code, with launch, regression, and headroom test layers. Software factory and agent collaboration (Priority: 9/5): They’re building multi-agent workflows for coding, debugging, review, and maintenance. MCP, CLIs, and tool architecture (Priority: 8/5): CLIs are favored for autonomy; MCP is kept for lightweight, permissioned integrations. Pricing, credits, and model selection (Priority: 8/5): Usage-based credits abstract over heterogeneous costs and let users choose speed, quality, or cost. Meeting notes and system of record (Priority: 8/5): Meeting notes and search create a data flywheel that improves collaboration and agent context.

Key Arguments: Reasoning models unlocked agents, but the bigger shift was better harnesses and tool design. Notion wins by serving enterprise work as the system of record, not by hiding that it uses wrappers. Evals are not one thing: launch checks, regression tests, and 30% headroom evals serve different purposes. Most failures come from tool bugs, so fixing tools beats retraining models in most cases. CLIs are great for self-debugging agents; MCP is best for narrow, tightly permissioned tools. Agents should use code where possible because deterministic execution is cheaper and more reliable. Pricing in credits was chosen because token costs vary by model, search, and sandbox infrastructure. The product should optimize user journeys like email triage and PDF export, not chase cool tools. Meeting notes create valuable organizational memory and increase both search quality and agent usefulness.

Data Points: team size (core AI capabilities and infrastructure): about 50 people - Engineering team leading core AI capabilities partner teams for packaging: another 30, 40 people - Teams that ship custom agents, meeting notes, and other surfaces custom agent launch performance: most successful launch in terms of free trials and converting people - Launch results after making custom agents free for three months system prompt/tooling scale: over a hundred tools - Latest version of the agent, not yet in custom agents tool ownership culture: five or six people - Earlier center-of-excellence era when few people could touch the few-shot prompt file manager-agent overload example: over 30 custom agents - A GTM team built many agents and needed a manager agent notifications reduced: from 70 per day to five - Manager agent reduced blocked-agent notifications eval pass rate target: 80 or 90 percent - Launch-quality category threshold for user journeys headroom eval pass rate: 30 percent - Frontier evals intentionally kept low to reveal future capability gaps deprecated model references: Sonnet 4, 3.7, 5.1, 5.2, 5.4 - Conversation about model churn and price differences price delta: 5.4 is 40% more expensive than 5.2 - Used to illustrate model pricing differences context length example: 17 days - A single coding-agent thread ran that long before being discovered as a harness bug launch delay: a few weeks - Custom agents launch was delayed while flipping the UX to chat-first delay estimate: a month or so - More specific estimate of the launch delay caused by the UX change

Pivotal Quotes: "MCP is just the dumb, simple thing that works, and it's pretty good." — Simon: Why Notion still supports MCP for narrow, tightly permissioned agents "Teach to the top of the class." — Sarah: Product philosophy: build deep tools for power users, not over-simplify "Notion is dedicated to being the best system of record for where people do their enterprise work." — Simon: Explaining why Notion supports MCP and builds its own integrations

Implications: Notion is betting that the next product moat is a tightly governed agent platform; the open question is how far general-purpose agent workflows can scale without losing quality or control.

🔓 Sign Up for Unlimited Episode Search

About Latent Space: The AI Engineer Podcast

The podcast by and for AI Engineers! In 2025, over 10 million readers and listeners came to Latent Space to hear about news, papers and interviews in Software 3.0. We cover Foundation Models changing every domain in Code Generation, Multimodality, AI Agents, GPU Infra and more, directly from the founders, builders, and thinkers involved in pushing the cutting edge. Striving to give you both the definitive take on the Current Thing down to the first introduction to the tech you'll be using in the next 3 months! We break news and exclusive interviews from OpenAI, Anthropic, Gemini, Meta (Soumith Chintala), Sierra (Bret Taylor), tiny (George Hotz), Databricks/MosaicML (Jon Frankle), Modular (Chris Lattner), Answer.ai (Jeremy Howard), et al. Full show notes always on https://latent.space

View all episodes from Latent Space: The AI Engineer Podcast