Lenny's Podcast
Lenny's Podcast

Anthropic’s first technical PM on token maxing, the jagged edge, and living in the future | Dianne Penn

Dianne Penn is Head of Product for Anthropic’s AI Research and Labs teams. She joined in 2023 as Anthropic’s first technical product manager, when the entire product team was five engineers, and has since helped ship every model from Claude 2 through Fable, and helped incubate Claude Code, MCP, Skil

Featured Speakers

Lenny Rachitsky HostDiane Penn Guest

Topics Discussed

Episode Summary

Executive Summary: Diane Penn, Anthropic’s head of product for research and labs, explains how the company evolved from a tiny startup into a frontier AI leader by pairing model breakthroughs with product experiences like Claude Code. She argues that evals are now central to product work, that ambitious, hands-on experimentation is essential, and that human judgment, taste, and first principles thinking remain critical in an AI-first world.

Main Topics: Anthropic’s early days and identity formation (Priority: 5/5): Diane describes Anthropic’s startup-like culture, tiny team, and early experiments that helped the company find its product identity, including quirky research demos that turned into public-facing experiences. Model/product co-evolution: Opus 3, Opus 4.5, and Claude Code (Priority: 5/5): The conversation highlights two major inflection points: Opus 3 as Anthropic’s frontier-model confidence boost, and Opus 4.5 plus Claude Code as the moment model capability and product experience reinforced each other. Evals as the new PRDs (Priority: 5/5): Diane argues that product teams in AI need to translate messy user feedback into concrete evals, which become the primary mechanism for defining problems, measuring progress, and aligning research and product. Labs as a system for discontinuous bets (Priority: 4/5): Anthropic’s labs org exists to explore high-upside ideas outside the core roadmap, using small, self-driven pods and a strong culture of experimentation to find 10x-to-1000x opportunities. What makes great AI researchers and PMs (Priority: 4/5): Successful researchers and PMs are first-principles thinkers who stay close to training runs, user pain points, and model behavior, while also being ambitious, hands-on, and comfortable with ambiguity. Safety, safeguards, and model access (Priority: 4/5): As models become more capable, Anthropic must strengthen safeguards, red teaming, and fallback UX while trying to keep powerful general-purpose AI broadly accessible. Human skills that still matter in the AI era (Priority: 4/5): Judgment, persistence, proactivity, curiosity, and an individual point of view remain highly valuable, especially as models get better at execution but still need human direction and verification.

Key Arguments: AI product work now requires adapting to emergent capabilities rather than following fixed plans; teams must be ready for what each new model can suddenly do. Evals are the operational bridge between vague user feedback and actionable research improvements, making them more useful than traditional PRDs in many AI workflows. Model capability and product experience are mutually reinforcing: frontier models need frontier products, and frontier products accelerate adoption of frontier models. High-performing teams in AI are small, hands-on, and deeply collaborative; experimentation works best when people work in public and learn together. Researchers succeed when they combine bold vision with close attention to details like training runs, evals, and data quality. AI should augment thinking, not replace it; the best use of Claude is as a sparring partner that improves judgment rather than blindly agreeing. Human judgment remains essential because current systems still lack lived experience and context, especially for deciding what to build and how to interpret nuanced user feedback.

Data Points: Anthropic early product team size: 5 engineers - Diane joined in 2023 when the product team was extremely small. Initial API staffing: 1 engineer - She noted there was only one engineer working on the entire API business early on. Anthropic company size during Opus 3: Less than 200 people - She described the Opus 3 launch as a company-wide effort at a still-small scale. Golden Gate Cloud demo reach: ~2,000 people - The quirky Golden Gate Bridge-themed Claude experience was spun up quickly and saw limited but meaningful reach. Claude code/token spending benchmark: $100,000/year - Referenced Gary Tan’s idea that people spending this much on tokens are living like the future. Model launch cadence in 2024: 4 model series in the whole year - Diane used this to illustrate the pace of progress before comparing it to 2025. Model launch cadence in 2025 Q2: More than 4 model series - She said Anthropic shipped more in one quarter than it had in all of 2024. Early eval set size for JSON instruction following: 30 to 40 examples - Used to build a first eval for Claude’s early schema-following failures. Instruction-following failure rate in early feedback: ~80% - Most of the user complaints were actually about Claude failing to write the right JSON. Current eval performance on that issue: 99.9% to 100% - She said the JSON-following issue is now essentially resolved on current models. Personal AI tenure: 6 years - She mentioned six years working in AI across Amazon and Anthropic.

Pivotal Quotes: "Evals are the new PRDs." — Diane Penn: Her core product philosophy for AI product management: translate user pain into measurable evals rather than relying only on traditional PRDs. "You have to sweat the tokens as much as you sweat the pixels." — Diane Penn: She was describing how AI product teams must pay attention to model behavior, usage, and experimentation, not just UI. "No matter how far you go, there’s always another level." — Diane Penn: A life motto from her grandfather that she uses to frame ambition, learning, and constant growth.

Implications: AI product teams will increasingly operate like research labs: hands-on, eval-driven, and highly adaptive. Human judgment, ambition, and taste become more important—not less—as models get more capable.

🔓 Sign Up for Unlimited Episode Search

About Lenny's Podcast

Lenny Rachitsky interviews world-class product leaders and growth experts about building products and growing careers.

View all episodes from Lenny's Podcast