Y Combinator Startup Podcast
Y Combinator Startup Podcast

Boris Cherny: Building Claude Code

Fresh off the launch of Opus 5, Claude Code creator Boris Cherny joins Diana Hu at Startup School 2026 to talk about what the newest models can do, how Claude Code came to be, and what it means to build products when the underlying capabilities keep accelerating.

Featured Speakers

Y Combinator Host

Topics Discussed

Episode Summary

Executive Summary: The conversation centers on Opus V/Claude 5 and how rapidly improving models are changing agentic software development. Boris argues that builders should aggressively delete outdated prompts and scaffolding, rely on empirical ablation, and let models run longer tasks with verification. He highlights new capabilities like resistance to prompt injection, long-running autonomous work, dynamic workflows, and routine-based self-maintenance, framing the future as one where developers “unhobble” models instead of over-constraining them.

Main Topics: Opus V’s new capabilities (Priority: 5/5): Boris describes major leaps in capability: longer autonomous runs, improved alignment, stronger vision/computer-use, and a lack of prompt-injection susceptibility compared with prior models. Delete-and-rebuild product harnesses (Priority: 5/5): He explains that system prompts, tools, and even harness code should be deleted or heavily reduced for each new model release, then rebuilt empirically based on observed failures. Prompt injection and security (Priority: 4/5): The discussion covers how layered defenses—alignment, classifiers, and auto-mode checks—make prompt injection much harder to demonstrate and change agent/product security design. Unhobbling models and product overhang (Priority: 5/5): Boris argues many products fail to elicit latent model abilities. The job of founders is to remove unnecessary scaffolding and create products that let models express hidden capabilities. Dynamic workflows, loops, and routines (Priority: 5/5): He presents new orchestration patterns that let Claude spawn and manage thousands of agents across sequential and parallel workflows for complex or repetitive tasks. Coding is increasingly solved, but not everywhere (Priority: 4/5): While agentic coding handles more everyday software work, some domains like distributed systems, deep systems code, and pixel-perfect UI verification still remain challenging. What students and builders should learn (Priority: 3/5): Boris advises students to learn practical application—building products, talking to users, and developing business/design sense—rather than only theoretical CS.

Key Arguments: New models can learn capabilities that were not explicitly taught, including sustained long-running work and improved resistance to prompt injection. System prompts and scaffolding should be treated as temporary; delete most of them when a new model arrives and only add back what repeated failures prove is necessary. A scientific, empirical workflow—try, observe, ablate, iterate—is now more useful than designing a large static system up front. Evals matter, but they also age quickly because models improve so fast that benchmarks get saturated and must be replaced. The biggest opportunity in agentic products is model elicitation: finding tasks and interfaces that let the model do what it is already capable of doing. Verification is the most important underused skill; give models ways to check their own work so they can safely tackle harder tasks. Dynamic workflows and routines turn test-time compute into a scalable orchestration layer, enabling thousands of agents to work productively. Founders should stop over-specifying tasks; modern models often perform better when given higher-level objectives, guardrails, and exit criteria instead of step-by-step instructions. The best users of Claude/ClotCode are those who are willing to experiment, forget old assumptions, and let the model surprise them. Practical skill-building—building for yourself, then for users—still matters even in an AI-heavy era.

Data Points: Arc AGI score: 3% to 30% - Boris says Opus V took Arc AGI from the previous best low-single-digit/low-teens range up to 30%. System prompt removed: 80%+ - He says the new release allowed ClotCode to delete more than 80% of its system prompt. Autonomous run duration: days, weeks, months - With auto mode, Opus 5 can run for extremely long periods without stopping. Bun rewrite duration: 11 days - A Bun/Zig codebase rewrite to Rust ran for 11 days using Claude and dynamic workflows. Desktop rewrite duration: 14-15 days - Boris says an ongoing task to rewrite the Electron desktop app in Swift had been running for a little over two weeks. Agents spawned: thousands, tens of thousands - He estimates large dynamic workflows can orchestrate thousands or even tens of thousands of agents. Routine frequency: daily / every 5 minutes / every hour - He describes routines and loops that run on schedules to maintain codebases and ship fixes. Models as of prior comparison: Sonnet 3.5 - He uses Sonnet 3.5 as the historical example of the first great coding model behind ClotCode’s original design. Internal maintenance scope: CLI, iOS, Android, desktop apps - Quad/ClotCode routines now maintain multiple product codebases automatically. Manual code share: more than 50% / 100% - Audience polling asks how much code is now written by agents rather than by hand.

Pivotal Quotes: "the model does not seem to be prompt-injectable anymore" — Boris: He describes Opus V’s improved resistance to prompt injection as a major security frontier. "press delete for everything" — Boris: His advice for rebuilding harnesses and prompts when new models ship. "coding is solved for the kind of coding that I do" — Boris: He qualifies the claim while arguing that agentic coding is already handling a large share of everyday development.

Implications: Agentic software is shifting from careful prompt engineering to empirical workflow design, verification, and orchestration. Builders who adapt quickly can unlock major leverage, while stale prompts, rigid harnesses, and over-engineering will increasingly hurt performance.

🔓 Sign Up for Unlimited Episode Search

About Y Combinator Startup Podcast

We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.

View all episodes from Y Combinator Startup Podcast