Episode Summary
Executive Summary: Francois Chollet argues that AI progress is irreversible and that the next frontier is not just scaling LLMs but building more optimal, symbolic, and agentic systems. He explains Endia’s program-synthesis approach, defends ARC as a benchmark for fluid intelligence, and says ARC v3 will test interactive, human-like exploration in unfamiliar environments. He predicts AGI around 2030 and urges builders to ride the wave rather than resist it.
Main Topics: Endia’s symbolic program-synthesis approach (Priority: 5/5): Chollet describes Endia as an AGI lab aiming to replace parametric deep learning with symbolic models learned via a new mechanism he calls symbolic descent, with the goal of producing concise, efficient, generalizable models. Why alternative AI approaches matter (Priority: 5/5): He argues the industry is over-concentrated on the LLM stack and that future AI should trend toward optimality, not just scale. He believes low-probability, high-upside research bets are worth pursuing if no one else is doing them. Verifiable rewards and the rise of coding agents (Priority: 5/5): Chollet says coding agents succeeded because code offers a formal verification signal, enabling RL loops that generate dense training data. He expects similar breakthroughs in other verifiable domains like mathematics. ARC as a benchmark for intelligence (Priority: 5/5): He explains ARC as a barometer for progress in fluid intelligence and reasoning, showing that base LLM scaling alone was insufficient while reasoning models and post-training loops caused major jumps. ARC v3 and agentic intelligence (Priority: 5/5): ARC v3 shifts from passive pattern modeling to interactive environments where models must explore, infer goals, and plan efficiently from scratch, aiming to measure human-like action efficiency. AGI definition and timeline (Priority: 4/5): Chollet distinguishes between automating economically valuable work and true general intelligence. He predicts AGI around 2030/early 2030s and says future systems may be much smaller and more elegant than current stacks. Advice for builders and open-source maintainers (Priority: 3/5): He emphasizes usability, community, and hiring power users, drawing on lessons from Keras and encouraging young people to use AI as leverage rather than fear job displacement.
Key Arguments: AGI should be defined as human-level learning efficiency across new tasks, not merely automation of economically valuable work. Current LLM-based systems are powerful but not the final architecture; AI research should move toward more optimal, symbolic foundations. Code and other formally verifiable domains are easier to automate because reward signals are trustworthy and can drive scalable RL loops. ARC v1 and v2 demonstrated that scaling pretraining alone was insufficient; reasoning and post-training were the real breakthroughs. ARC v3 is designed to be harder to game because it tests interactive exploration and goal discovery in unseen environments. A successful AGI system may ultimately be small in code size and model footprint, with a large knowledge base layered beneath it. Research directions with low odds but high upside are worth pursuing if they are neglected and could create a new branch of machine learning. Human involvement should be minimized in the improvement loop; scalable systems should improve themselves through data, search, and verification.
Data Points: ARC v1 performance of base models: sub-10% - Chollet says base LLMs remained extremely weak on ARC v1 even after massive scaling. Scaling of pretraining: 50,000x - He notes that base-model performance on ARC v1 stayed low despite roughly 50,000x scale-up. Chance of Endia success: 10–15% - Chollet estimates Endia’s approach has a low but worthwhile probability of success. G-Stack stars: 40,000 stars - The host references a viral open-source project milestone reached that morning. Contributor pull requests: 100+ PRs - The host mentions the open-source project now has over 100 pull requests to handle. ARC task cost: $0.3 cents per task - The host contrasts ARC-style task costs with foundation-model costs. Foundation model task cost: $1 to $10 per task - The host cites higher per-task costs for LLM-based approaches on the same kind of task. ARC v3 game count: 250+ games - Chollet says the ARC v3 studio created over 250 interactive games. Human solve time per ARC v3 game: ~10 minutes or less - He says each game is designed to be quickly playable from scratch. ARC v1 base-model score: zero for GPT-3 - He notes early GPT-3 scored zero on ARC v1. ARC v1 base-model score: extremely low, under 10% - He says even later base models remained very low on ARC v1. ARC v2 saturation: 97% - The host cites a company saturating ARC v2 with a 97% result. Keras release timing: March 2015 - Chollet says Keras was released exactly 11 years prior to the interview date. ARC paper publication: 2019 - He says the ARC paper and benchmark framing were published in 2019. AGI timeline: 2030 / early 2030s - Chollet predicts AGI around the time ARC 6 or ARC 7 might be released. ARC v3 private/public split: private set significantly different from public set - He says the public set is easier and not representative of private-set performance.
Pivotal Quotes: "I think we're probably looking at AGI 2030." — Francois Chollet: Opening prediction about the likely timeline for AGI. "AI progress is here. It's actually going to keep accelerating. How do you make use of it? How do you leverage? How do you ride the wave?" — Francois Chollet: His advice to builders and listeners about responding to AI progress. "Science is fundamentally a symbolic compression process." — Francois Chollet: He uses science as an analogy for Endia’s symbolic learning approach.
Implications: The industry may be underinvesting in non-LLM paradigms. For builders, the opportunity is to target verifiable domains, build compounding systems, and use AI as leverage rather than fear it. ARC v3 suggests the next benchmark frontier is agentic, not just predictive.
About Y Combinator Startup Podcast
We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.