Y Combinator Startup Podcast
Y Combinator Startup Podcast

François Chollet: The ARC Prize & How We Get to AGI

François Chollet on June 16, 2025 at AI Startup School in San Francisco. François Chollet is a leading voice in AI. He's the creator of the Keras library, author of Deep Learning with Python, and the founder of the ARC Prize, a global competition aimed at measuring true general intelligence. He

Featured Speakers

Y Combinator Host

Topics Discussed

Episode Summary

Executive Summary: Francois argues that scaling pretraining alone did not produce AGI/HGI because benchmarks rewarded memorized skills, not fluid intelligence. He proposes a new framework centered on information efficiency, abstraction, and test-time adaptation, using ARC as a diagnostic benchmark. The future, he says, lies in combining deep learning with discrete program search to build AI that can invent, adapt, and support scientific discovery.

Main Topics: Why pretraining scaling fell short (Priority: 5/5): The talk opens with the claim that compute has grown exponentially, enabling deep learning, but that larger models and more data mostly improved memorized task performance rather than genuine general intelligence. Defining intelligence as adaptation to novelty (Priority: 5/5): Francois contrasts task mastery with fluid intelligence, arguing that intelligence is the efficiency of using past experience to handle unfamiliar situations, not just performing known tasks. ARC as a benchmark for fluid intelligence (Priority: 5/5): ARC is presented as a benchmark designed to expose the gap between human-like reasoning and benchmark-optimized AI, first via ARC1 and then more sensitively via ARC2. Test-time adaptation as the new paradigm (Priority: 5/5): The speaker says 2024 marked a pivot from static pretraining to models that can modify their behavior at inference time, enabling progress on ARC and signaling a shift in AI methods. Two kinds of abstraction: type 1 and type 2 (Priority: 4/5): He distinguishes continuous/value-based abstraction (perception, intuition, transformers) from discrete/program-based abstraction (reasoning, planning, program search), arguing that human intelligence combines both. Program search plus deep learning as the next architecture (Priority: 5/5): The proposed path forward is a programmer-like meta-learner that uses deep learning for intuition and discrete search for synthesis, supported by a reusable library of abstractions. Building AI for scientific discovery (Priority: 4/5): The final goal is not just automation but autonomous invention and discovery, with the lab 'Endia' aiming to accelerate science through systems that can learn, adapt, and improve over time.

Key Arguments: Pretraining scaling produced large gains on benchmarks, but those benchmarks mainly measured static skills, not fluid intelligence. ARC1 showed that even massive scale-ups left performance near zero on tasks humans solve easily, implying pretraining alone is insufficient. Test-time adaptation is necessary for genuine fluid intelligence because it allows models to reprogram or modify themselves during inference. Benchmarks shape progress, so the field needs tests that measure novelty-handling, abstraction, and efficiency rather than memorization. ARC2 is a more sensitive test than ARC1 because it is harder to brute-force and better separates true reasoning from memorization. Intelligence should be measured by information efficiency, operational range, and the ability to recombine abstractions for new situations. Human cognition likely combines two abstraction modes: continuous pattern recognition and discrete program-like reasoning. Deep learning excels at type 1 abstraction but struggles with type 2 tasks such as exact symbolic manipulation and compositional reasoning. Discrete program search is necessary for invention and creativity, while deep learning provides the intuition needed to make search tractable. The future AI system will be a meta-learner that assembles task-specific programs from reusable building blocks and improves its own abstraction library over time.

Data Points: Compute cost decline: Two orders of magnitude every decade since 1940 - Used to explain why AI progress accelerated and why scaling became central to deep learning. ARC1 tasks: 1,000 tasks - Described as unique problems that cannot be crammed for and require on-the-fly reasoning. ARC1 performance increase: About 50,000x scale-up of basal models; accuracy from 0% to roughly 10% - Illustrates that massive scaling barely improved performance on the fluid-intelligence benchmark. Human performance on ARC1: Well above 95% - Used to highlight the gap between humans and frontier models on fluid reasoning. ARC2 human study: 400 people tested in person - The benchmark was validated using in-person human testing in San Diego. ARC2 task solvability: At least 2 people solved each task; about 7 people saw each task on average - Supports the claim that ARC2 remains easy for humans but difficult for AI. ARC2 crowd performance: 10 random people with majority voting would score 100% - Demonstrates ARC2 is feasible for ordinary humans without prior training. Bayes/LLM performance on ARC2: 0% - Static pretrained models like GPT-4.5 and Llama 4 reportedly fail entirely on ARC2. Static reasoning systems on ARC2: About 1% to 2% - Single-chain reasoning methods remain near zero, reinforcing the need for test-time adaptation. Human abstraction/data efficiency gap: Gradient descent needs roughly 3 to 4 orders of magnitude more data than humans - Used to argue current learning methods are far less sample-efficient than human cognition. ARC3 release timeline: Developer preview in July; launch in early 2026 - The next benchmark will assess agency, interactive learning, and goal-directed action efficiency.

Pivotal Quotes: "Intelligence is the efficiency with which you operationalize the past to face a constantly changing future." — Francois: Core definition of intelligence offered early in the talk. "You hit the targets, but you miss the points." — Francois: Explains how benchmark optimization can produce systems that look successful without advancing true intelligence. "Deep learning doesn't invent, but search does." — Francois: Summarizes the argument that invention requires discrete search, not just scalable pattern learning.

Implications: If this view is right, progress toward AGI/HGI will depend less on larger models and more on systems that can adapt, search, and recombine abstractions efficiently. That shifts industry focus toward benchmark design, agentic reasoning, and AI for scientific discovery.

🔓 Sign Up for Unlimited Episode Search

About Y Combinator Startup Podcast

We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.

View all episodes from Y Combinator Startup Podcast