Sean Carroll MindScape
Sean Carroll MindScape

280 | François Chollet on Deep Learning and the Meaning of Intelligence

Which is more intelligent, ChatGPT or a 3-year old? Of course this depends on what we mean by "intelligence." A modern LLM is certainly able to answer all sorts of questions that require knowledge far past the capacity of a 3-year old, and even to perform synthetic tasks that seem remarkab

Featured Speakers

Sean Carroll | Wondery HostFrancois Chollet Guest

Topics Discussed

Episode Summary

Executive Summary: Sean Carroll interviews François Chollet about why current large language models are powerful pattern-matchers but not generally intelligent. Chollet argues intelligence means robust adaptation to novelty, not just high performance on familiar tests. He discusses ARC as a benchmark for novel reasoning, critiques AGI timelines based on scaling, and emphasizes that LLMs are useful tools, but not yet systems that reason, plan, or invent like humans.

Main Topics: Symbolic AI vs. machine learning and deep learning (Priority: 4/5): Chollet distinguishes hand-coded symbolic systems from learned systems, noting that early machine learning successes came from non-neural methods like SVMs, random forests, and gradient boosting before deep learning resurged. Why LLMs appear intelligent (Priority: 5/5): LLMs can memorize and interpolate among patterns in vast training data, which lets them mimic human-like language, solve familiar tasks, and even combine styles such as 'Shakespearean pirate.' Limits of LLM reasoning, planning, and generalization (Priority: 5/5): Chollet argues LLMs fail on unfamiliar, slightly modified, or compositional tasks because they retrieve learned programs rather than synthesize new ones on the fly. Intelligence as adaptation to novelty (Priority: 5/5): He defines intelligence as the ability to pick up new skills and solve genuinely new problems efficiently from few examples, contrasting this with static model inference and memorization. ARC benchmark and competition design (Priority: 4/5): The Abstraction and Reasoning Corpus (ARC) was created to test human-like generalization on novel puzzles that are not in training data; it is now tied to a competition incentivizing new AI approaches. Current and future uses of deep learning and Keras (Priority: 3/5): Despite skepticism about AGI, Chollet is enthusiastic about accessible deep learning tools, especially Keras, and sees practical value in fine-tuning models for domain-specific tasks. AGI risk and autonomy (Priority: 4/5): Chollet argues that intelligence alone is not existentially dangerous; risk requires deliberate engineering of autonomy, goal-setting, and value systems, not just smarter prediction models.

Key Arguments: LLMs are best understood as massive pattern/retrieval systems that memorize programs or solution templates, then reapply them; this explains their fluency without implying human-like understanding. True intelligence requires generalization to novel situations, not merely strong performance on tasks similar to training examples. Many benchmark victories reflect memorization and test familiarity rather than reasoning; novel-problem benchmarks like ARC are more discriminating. LLMs can sometimes be improved with pointwise patches or fine-tuning, but this is evidence of brittle, local fixes rather than deep understanding. In-context learning is mischaracterized as learning; Chollet says LLMs are mostly fetching an internal rule/template, not genuinely updating themselves. Humans, including young children, can learn from very few examples and invent new solutions; current LLMs cannot match this sample efficiency or creative abstraction. AGI is not on a smooth scaling trajectory from current LLMs; if it arrives, it will require new ideas, not just bigger models and more data. AI danger is often overstated by conflating intelligence with autonomous agency; a capable system becomes threatening mainly when deliberately given goals, autonomy, and power. Deep learning tools like Keras are valuable because they democratize model building and fine-tuning, letting individuals and businesses adapt models to specific needs.

Data Points: Keras users: 3 million+ - Carroll notes the software library’s broad adoption in deep learning. ARC public benchmark human performance: ~80% - Carroll describes ARC as easy for humans but hard for AI. ARC competition machine performance: 0% to 20-30% - LLMs and other methods perform poorly on the novel puzzle benchmark. Arcathon prize pool: over $1,000,000 - Chollet says a rebooted Kaggle competition will offer large prizes to solve ARC. ARC task count in competition submission: 100 hidden tasks - Competitors submit programs evaluated on hidden puzzles. Compute limit per submission: 12 hours - Competition entrants get limited runtime to solve hidden ARC tasks. Hardware limit: 1 P100 GPU + multi-core CPU - Chollet explains the allowed compute budget for submissions. Model size feasible under constraints: ~8 billion parameters - He says float16 open-source LMs of that scale could fit within the competition constraints. Human sample efficiency example: fewer than 1,000 Lego bricks - Chollet cites his 3-year-old’s ability to invent new Lego constructions from limited exposure. Education/benchmarking example: 10 exams vs. 10,000 exams - He contrasts human cramming with LLM-scale memorization to explain apparent competence.

Pivotal Quotes: "intelligence according to me is the ability to pick up new skills to adapt to new situations to things you've not seen before" — Francois Chollet: Chollet defines intelligence in contrast to memorization and benchmark performance. "LLMs are basically program databases" — Francois Chollet: He summarizes his view that models store and reapply patterns rather than reason like programmers. "intelligence is is pretty much just a conversion ratio between the information you have to the ability to operate in novel situations in the future" — Francois Chollet: He explains intelligence as transforming prior experience into flexible future action.

Implications: For listeners and industry, the takeaway is to treat LLMs as powerful but brittle tools, not substitutes for general intelligence. Progress toward AGI likely requires new architectures focused on novelty, reasoning, and compositional abstraction, while current models remain best used as assistive, domain-specific systems.

🔓 Sign Up for Unlimited Episode Search

About Sean Carroll MindScape

Ever wanted to know how music affects your brain, what quantum mechanics really is, or how black holes work? Do you wonder why you get emotional each time you see a certain movie, or how on earth video games are designed? Then you’ve come to the right place. Each week, Sean Carroll will host conversations with some of the most interesting thinkers in the world. From neuroscientists and engineers to authors and television producers, Sean and his guests talk about the biggest ideas in science, ...

View all episodes from Sean Carroll MindScape