Lenny's Podcast
Lenny's Podcast

The 100-person AI lab that became Anthropic and Google's secret weapon | Edwin Chen (Surge AI)

Edwin Chen is the founder and CEO of Surge AI, the company that teaches AI what’s good vs. what’s bad, powering frontier labs with elite data, environments, and evaluations. Surge surpassed $1 billion in revenue with under 100 employees last year, completely bootstrapped—the fastest company in histo

Featured Speakers

Lenny Rachitsky HostEdwin Chen Guest

Topics Discussed

Episode Summary

Executive Summary: Edwin Chen, founder/CEO of Surge AI, argues that high-quality human data—not just scale or benchmarks—is what truly advances AI. He explains how Surge bootstrapped to over $1B revenue with a tiny elite team, why Silicon Valley’s fundraising/PR machine can distort company building, and why current AI incentives may push models toward engagement and “AI slop” instead of truth, usefulness, and real-world capability.

Main Topics: Surge’s unusually efficient, bootstrapped scale (Priority: 5/5): Chen describes Surge’s rise to over $1B in revenue in under four years with under 100 people, no VC funding, and profitability from day one. He frames this as proof that small, elite teams can outperform large organizations, especially as AI increases leverage. Defining and producing high-quality AI data (Priority: 5/5): A central theme is that quality is nuanced and cannot be produced by simply hiring more people. Surge uses deep signals, expert judgment, and performance analysis to identify what 'good' means in domains like poetry, coding, and documentation, then uses that to teach models. Why benchmarks and leaderboards can mislead (Priority: 5/5): Chen says many benchmarks are flawed, easy to game, or too objective compared with messy real-world tasks. He criticizes practices like leaderboard chasing and argues that benchmark optimization can distort actual model capability and public perception. AI values shape model behavior (Priority: 4/5): He argues that labs’ values and objective functions determine model personality—whether a model optimizes for productivity, engagement, flattery, or truth. This leads to differentiated models over time rather than a single commoditized standard. Reinforcement learning environments as the next frontier (Priority: 5/5): Chen sees RL environments as a major next step: simulated real-world tasks with tools, long horizons, and rewards based on outcomes and trajectories. He believes these will better reflect how humans learn and expose model weaknesses that benchmarks miss. A critique of Silicon Valley startup norms (Priority: 4/5): Chen rejects common advice to pivot constantly, blitzscale, raise money early, and perform for VCs. He argues founders should build only what they uniquely can, stay mission-focused, and avoid the fundraising/PR hamster wheel. Surge as a research lab, not just a startup (Priority: 4/5): Surge invests heavily in internal research to build better benchmarks, improve training data, and understand model behavior. Chen says his motivation is scientific curiosity and shaping AI’s direction, not just chasing revenue or valuation.

Key Arguments: Small, elite teams can create extreme leverage; AI will make per-employee revenue ratios even larger over time. Quality in AI data is subjective, domain-specific, and requires sophisticated human judgment and many signals—not just more annotators. Benchmarks often measure what is easy to score or game, not what matters in the real world. Frontier labs’ values and incentives directly shape model behavior and outputs. Current incentives may optimize models for engagement, flattery, and superficial polish rather than truth or usefulness. RL environments can train models on messy, end-to-end tasks closer to real work than static benchmarks. The best founders should build what only they can build rather than chase conventional VC playbooks. A company’s objective function matters: good AI should make users more capable, curious, and creative, not merely more engaged.

Data Points: Revenue: Over $1 billion - Surge reportedly hit this level last year. Employee count: Under 100 people - Chen repeatedly emphasizes Surge’s tiny team relative to its scale. Time to $1B revenue: Less than four years - A key measure of Surge’s growth speed. Funding: $0 VC raised - Surge is completely bootstrapped. Profitability: Profitable from day one - Chen notes Surge never needed outside capital to survive. Potential future revenue per employee: $100 million per employee - Chen predicts even higher efficiency ratios in coming years. Model job coverage estimate: 80% of an average L6 software engineer’s job in 1-2 years - Chen’s near-term automation outlook. Longer-term progression: 90% in a few more years; 99% and 99.9% later - He argues model improvement will be incremental and non-linear near the top end. Annotation feedback scale: Thousands of signals - Surge tracks extensive signals on workers, tasks, and projects to assess quality. Poem example: Eight-line poem about the moon - Used to explain the difference between checkbox compliance and true quality. Email drafting example: 30 minutes and 30 versions - Chen describes Claude helping craft an email but at a productivity cost. Leaderboard gaming example: More emojis, more markdown, longer responses - He says these tactics can improve appearance on LM Arena even if accuracy worsens. Coverage of company work: Every frontier AI lab - The intro states Surge powers training at every frontier AI lab.

Pivotal Quotes: "We basically never wanted to play the Silicon Valley game. I always thought it was ridiculous." — Edwin Chen: Explaining why Surge avoided the usual fundraising, PR, and networking playbook. "We are basically teaching our models to chase dopamine instead of truth." — Edwin Chen: His critique of current AI incentives, benchmarks, and engagement-driven behavior. "I would rather be Terrence Tao than Warren Buffett." — Edwin Chen: Describing his identity as a scientist/researcher first and a business person second.

Implications: The conversation suggests AI progress will depend as much on incentive design, data quality, and human judgment as on model scale. For founders, it’s a case for mission-first, research-driven companies; for the industry, a warning that current benchmarks may reward the wrong behaviors.

🔓 Sign Up for Unlimited Episode Search

About Lenny's Podcast

Lenny Rachitsky interviews world-class product leaders and growth experts about building products and growing careers.

View all episodes from Lenny's Podcast