Episode Summary
Executive Summary: Mike Noop, co-founder and head of AI at Zapier, explains why he believes AGI progress has stalled, arguing that current LLM-based systems are useful but not true general intelligence. He introduces Ark Prize, a nonprofit challenge centered on the ARC AGI eval, to incentivize outsiders to discover open-source breakthroughs in efficient skill acquisition, program synthesis, and novel AI architectures.
Main Topics: Why AGI progress has stalled (Priority: 5/5): Noop argues the field has over-indexed on LLM scaling while missing the core property of general intelligence: efficient acquisition of new skills. He says the consensus AGI definition is flawed and frontier labs have become less transparent. Ark Prize and ARC AGI eval (Priority: 5/5): He describes Ark Prize as a million-dollar-plus nonprofit public challenge to beat ARC AGI, open-source the solution, and spur progress toward a more accurate AGI benchmark. What AGI should mean (Priority: 5/5): Noop contrasts the common 'economically useful work' definition with Francois Chollet's definition: a system that can efficiently acquire new skills and solve open-ended problems. He favors the latter as a better measure of true intelligence. Limits of LLM scaling (Priority: 4/5): He argues that LLMs are high-dimensional memorization systems that can generalize patterns but cannot invent or discover where no training-data pattern exists. He believes scaling alone will not produce AGI. Promising technical directions (Priority: 4/5): Noop highlights program synthesis and neural architecture search as underexplored paths that may better support general intelligence, especially if guided by cheap compute and less human bias. Zapier's AI product evolution (Priority: 4/5): He describes how Zapier explored chain-of-thought, tree-of-thought, and ChatGPT-like systems internally, then found tool use and workflow automation to be the clearest product direction, leading to 50M+ AI tasks and AI bots. Open source and regulation (Priority: 4/5): Noop strongly favors open research and open source for foundational AI progress, while arguing that current AI should be governed via existing regulatory frameworks rather than prescriptive AGI legislation based on speculation.
Key Arguments: The dominant AGI metric is wrong because measuring 'economic usefulness' says more about human labor markets than about intelligence itself. True general intelligence is the efficient acquisition of new skills; that is the capability humans display and current models largely lack. LLMs can memorize and generalize from patterns in training data, but they cannot reliably solve problems whose solution patterns are absent from the data. ARC AGI is, in Noop's view, the only real AGI eval because it resists scale and LMs and has shown only modest progress over four years. Outsiders and small teams are more likely to make the breakthrough than incumbent labs, so a prize structure is better than traditional startup investment. Program synthesis is promising because it searches a broad program space rather than optimizing within a narrow learned representation. Neural architecture search should be revisited now that compute is cheaper, with less human bias and more relaxed search. Open source accelerates foundational discovery by exposing more researchers to ideas and counteracting the closure of frontier lab publishing. AI should be regulated through existing frameworks, while prescriptive AGI restrictions are premature without empirical evidence of capabilities.
Data Points: Ark Prize funding: $1 million+ - Nonprofit public challenge to beat ARC AGI and open-source the solution ARC AGI state of the art: 34% - Best reported performance today on the eval ARC AGI state of the art at launch: 20% - Performance when the eval was introduced four years earlier ARC challenge team count: 300 teams - Number of teams competing in the smaller 2023 contest Zapier AI tasks run: 50 million+ - AI tasks executed on Zapier over the last year and a half Zapier integrations: 6,000 - Number of integrations Zapier can use with AI workflows ARC puzzle board size: 2x2 to 15x15 squares - Noop describes the eval as small and reproducible Estimated solution size: 10,000 lines of code or less - Noop suggests the ARC solution may be compact rather than requiring giant models Timeframe of AGI stall: 4-5 years - Noop says AGI progress has stalled over this period Time spent exploring AI at Zapier: about 6 months - He and the CTO went back to IC roles to test what was possible
Pivotal Quotes: "I think we're measuring the wrong things." — Mike Noop: On why the consensus definition of AGI is flawed "General intelligence is a system that can effectively, efficiently acquire new skill." — Mike Noop: On Francois Chollet's definition of AGI "We need something more. There's something additional in addition to that we need." — Mike Noop: On why scaling LLMs alone will not reach AGI
Implications: The conversation frames AGI as a skills-acquisition problem, not just a scaling problem. It suggests future breakthroughs may come from outsiders, open research, and hybrid approaches beyond LLMs, with implications for product strategy, funding, and regulation.