Lenny's Podcast
Lenny's Podcast

“Engineers are becoming sorcerers” | The future of software development with OpenAI’s Sherwin Wu

Sherwin Wu leads engineering for OpenAI’s API platform, where roughly 95% of engineers use Codex, often working with fleets of 10 to 20 parallel AI agents. We discuss: 1. What OpenAI did to cut code review times from 10-15 minutes to 2-3 minutes 2. How AI is changing the role of managers 3. Why the

Featured Speakers

Lenny Rachitsky HostSherwin Wu Guest

Topics Discussed

Episode Summary

Executive Summary: Sherwin Wu, head of engineering for OpenAI’s API and developer platform, argues that AI is rapidly transforming software engineering from writing code to managing agents, and that the biggest near-term opportunities may be in business process automation and AI-enabled B2B SaaS. He emphasizes building for where models are going, not where they are today, and says the winners will pair top-down support with bottoms-up adoption.

Main Topics: AI is changing the engineer’s job (Priority: 5/5): Engineers at OpenAI increasingly use Codex for writing and reviewing code, shifting the role from individual contributor to agent manager and technical lead. Managing agents and workflow stress (Priority: 4/5): As agentic coding becomes normal, teams experience stress when agents fail; OpenAI is learning best practices by running a fully Codex-written codebase experiment. Management in an AI world (Priority: 4/5): Managers will likely become higher-leverage operators who spend more time with top performers, unblock work, and use AI to understand organizational context and performance. Second-order startup effects (Priority: 5/5): Wu predicts a one-person billion-dollar startup could spawn a huge ecosystem of niche software startups, potentially ushering in a golden age of B2B SaaS and changing venture dynamics. Building AI products with model trajectory in mind (Priority: 5/5): He warns that models quickly obsolete scaffolding and recommends building for future model capabilities rather than current limitations. Enterprise AI adoption and ROI (Priority: 4/5): Many deployments underperform because they are forced top-down without bottom-up employee buy-in; successful adoption requires internal champions and practical workflows. OpenAI platform strategy and ecosystem (Priority: 4/5): OpenAI sees itself as a platform company, aiming to spread AI benefits broadly through the API, agents tools, and ecosystem support rather than squashing startups.

Key Arguments: AI is shifting engineering from code production to agent orchestration; high-performing engineers now manage many parallel Codex threads. Code review and CI are being automated, making review faster and reducing friction before deployment. The hardest AI failures often come from missing context, so teams need better documentation and encoded tribal knowledge. Managers will become more leveraged by using AI to surface blockers, summarize org knowledge, and scale team oversight. The biggest AI opportunity may be outside software engineering—in repeatable business processes across enterprises. Top-down AI mandates fail without bottom-up enthusiasm and local experts who can teach and evangelize usage. Startups should not fear OpenAI crushing them; the market is large enough and value comes from building something users genuinely love. Successful AI products should target an ideal capability that models are almost ready for, then ride model improvements as they arrive.

Data Points: OpenAI engineers using Codex daily: 95% - Wu says nearly all engineers at OpenAI use Codex every day. PRs reviewed by Codex: 100% - All pull requests at OpenAI are reviewed by Codex. Code written/generated by AI first: close to 100% - Wu says the vast majority of code is AI-generated before human review. PR output difference for Codex users: 70% more PRs - Engineers who use Codex more tend to open substantially more pull requests. Agent threads handled by engineers: 10 to 20 threads - Some engineers simultaneously manage many Codex/agent tasks. Top performer management focus: more than 50% of time - Wu says he spends over half his time with top performers. ChatGPT weekly active users: 800 million - Wu cites the scale of ChatGPT’s user base. AI task duration benchmark: multi-hour tasks at 50% success; just under 1 hour at 80% - He references the meter benchmark for coherent software-engineering task length.

Pivotal Quotes: "This is the worst the models will ever be." — Kevin Weil (quoted by Sherwin Wu): Used to illustrate why model capabilities will keep improving and why current limitations are temporary. "The models will eat your scaffolding for breakfast." — Nicholas (quoted by Sherwin Wu): Used to explain how product scaffolding and agent frameworks can become obsolete as models improve. "Make sure you’re building for where the models are going and not where they are today." — Sherwin Wu: Wu’s core advice to builders creating products on top of AI models and APIs.

Implications: Expect engineers to become agent managers, managers to become more leveraged, and AI products to shift toward longer-running tasks and audio. The biggest opportunities may be in business automation and niche SaaS built for a post-AI workflow world.

🔓 Sign Up for Unlimited Episode Search

About Lenny's Podcast

Lenny Rachitsky interviews world-class product leaders and growth experts about building products and growing careers.

View all episodes from Lenny's Podcast