Episode Summary
Executive Summary: Greg Camrat explains ArcPrize’s mission to advance open AGI research through benchmarks that measure how well systems learn new things, not just how hard they can test. He contrasts static benchmarks like ARC-AGI 1 and 2 with the upcoming interactive ARC-AGI 3, which will assess generalization, sample efficiency, and human-like problem solving in game-like environments.
Main Topics: ArcPrize’s mission and definition of intelligence (Priority: 5/5): Camrat frames ArcPrize as a tech-forward nonprofit focused on open progress toward systems that generalize like humans. Intelligence is defined as the ability to learn new things efficiently, following François Chollet’s theory. Why ARC differs from standard benchmarks (Priority: 5/5): ARC is designed to measure learning and generalization on tasks regular people can solve, rather than pushing toward ever-harder expert-level tests like MMLU or Humanity’s Last Exam. Benchmark results and the reasoning paradigm (Priority: 5/5): Camrat notes that pre-reasoning base models performed poorly on ARC, but reasoning models dramatically improved performance, helping reveal the importance of reasoning in frontier AI. Adoption by frontier labs (Priority: 4/5): OpenAI, xAI, Google Gemini, and Anthropic now report ARC-related performance, which validates the benchmark but also raises concerns about vanity metrics overshadowing the broader mission. False positives in AI progress (Priority: 4/5): Camrat criticizes the idea that building custom RL environments for every domain equals general intelligence, arguing this is a short-term workaround rather than real generalization. ARC-AGI 3 and interactive evaluation (Priority: 5/5): The next version will be a hidden, interactive benchmark with about 150 video-game-like environments that require action-feedback learning without instructions, closer to real-world intelligence. Beyond accuracy: efficiency, data, and energy (Priority: 5/5): ARC-AGI 3 will evaluate not only correctness but also how many actions, how much data, and how much energy a system needs versus humans to solve tasks.
Key Arguments: Intelligence should be measured by the ability to learn novel tasks efficiently, not just by static test scores. ARC benchmarks are meaningful because ordinary humans can solve them, making human-vs-AI comparison more informative. ARC-AGI exposed the limits of base models and highlighted the importance of reasoning models. Frontier lab adoption is useful for visibility, but benchmark reporting can become a vanity metric if it does not advance the core mission. Creating one-off RL environments for each problem is a whack-a-mole approach and does not scale to real-world novelty. ARC-AGI 3 will be more realistic because it introduces interaction, feedback, and hidden objectives rather than explicit instructions. Future AGI evaluation should include sample efficiency and energy efficiency, since those better reflect how humans learn and operate.
Data Points: ARC-AGI 1 launch year: 2019 - François Chollet introduced the original ARC benchmark and the measure-of-intelligence framing. ARC-AGI 2 launch year: 2025 (March) - Camrat says ARC-AGI 2 was released earlier in the year as a deeper static benchmark. ARC-AGI 3 release timing: Next year - The interactive third version is planned for the following year. ARC-AGI 3 environments: ~150 - Camrat says V3 will consist of about 150 video-game-like environments. GPT-4 base model ARC score: 4–5% - He cites base GPT-4 as performing very poorly on ARC before reasoning models. o1 / o1 preview ARC score: 21% - He says early reasoning models substantially improved ARC performance. Human participants per ARC-AGI 3 game: 10 people - Each environment will be validated by recruiting ordinary people to test solvability. Human solvability requirement: Minimum threshold - If a game fails the threshold among regular humans, it will be excluded. Frontier lab adopters mentioned: OpenAI, xAI, Gemini, Anthropic - Camrat lists major labs using ARC for model release reporting. ARC-AGI 1 task count: 800 tasks - He notes Chollet reportedly made all 800 tasks himself for the first benchmark.
Pivotal Quotes: "intelligence as your ability to learn new things" — Greg Camrat: Explaining the benchmark philosophy derived from François Chollet’s 2019 paper. "the thing that solves Arc AGI is necessary for AGI, it's not sufficient" — Greg Camrat: Clarifying that benchmark success indicates strong generalization but does not by itself equal AGI. "we're not going to give any instructions to the test taker on how to complete the environment" — Greg Camrat: Describing the design of ARC-AGI 3 as an interactive, instruction-free evaluation.
Implications: ARC pushes AI evaluation toward human-like learning, interaction, and efficiency. If adopted widely, it could shift focus from leaderboard-chasing to systems that generalize in novel, real-world settings and reveal what current models still lack.
About Y Combinator Startup Podcast
We help founders make something people want. The Y Combinator Podcast is where builders talk about building. From the earliest days of an idea to scaling a company that changes the world, YC partners and founders share real stories, lessons, and tactics from the frontlines.