The Twenty Minute VC (20VC)
The Twenty Minute VC (20VC)

20VC: AI Scaling Myths: More Compute is not the Answer | The Core Bottlenecks in AI Today: Data, Algorithms and Compute | The Future of Models: Open vs Closed, Small vs Large with Arvind Narayanan, Professor of Computer Science @ Princeton

Arvind Narayanan is a professor of Computer Science at Princeton and the director of the Center for Information Technology Policy. He is a co-author of the book AI Snake Oil and a big proponent of the AI scaling myths around the importance of just adding more compute. He is also the lead author of a

Featured Speakers

Arvind Narayanan Guest

Topics Discussed

Episode Summary

Executive Summary: Arvind Narayanan argues that AI progress is entering a new phase where data, not compute, is the main bottleneck, making ever-larger frontier models less likely to deliver GPT-4-style leaps. He expects smaller, cheaper, more deployable models, emphasizes product usefulness over AGI hype, and warns that many AI risks are better framed as existing social, policy, and platform problems rather than technology problems alone.

Main Topics: Compute, data, and diminishing returns (Priority: 5/5): Narayanan argues that scaling compute alone is yielding smaller gains because frontier models are already trained on most accessible data; future progress is more likely to come from smaller, more efficient models and new ideas beyond brute-force scaling. Synthetic data and hidden data sources (Priority: 5/5): He pushes back on claims that untapped sources like YouTube or synthetic data will unlock another era of scaling, saying video-to-text token yields are smaller than they seem and synthetic data mostly risks recycling existing knowledge rather than creating new capability. Model commoditization and inference economics (Priority: 4/5): The discussion highlights how cost pressures and deployment realities are pushing the market toward smaller models, on-device inference, and a world where the largest value may accrue above the base model layer. AGI hype vs product reality (Priority: 5/5): Narayanan criticizes AI companies for over-indexing on AGI and underinvesting in product design, arguing that real utility, product-market fit, and transparency matter more than treating chatbots as demos or proto-gods. Evaluation, benchmarks, and real-world usefulness (Priority: 4/5): He says LLM benchmarks are increasingly unreliable because of contamination and optimization for test performance rather than actual usefulness; real-world feedback from professionals matters more than leaderboard scores. Policy, regulation, and misinformation (Priority: 4/5): Narayanan frames much of ‘AI regulation’ as ordinary regulation of harmful activities enabled by AI, and argues that misinformation is mainly a distribution/social media problem rather than an AI generation problem. Societal impacts in education, medicine, jobs, and defense (Priority: 5/5): He is skeptical of sweeping claims about AI replacing doctors or most jobs, stressing that AI automates tasks not whole jobs, and that the biggest concerns may be deepfake nudes, education costs, and ensuring AI helps defense more than offense.

Key Arguments: More compute still helps, but the easy gains from simply making frontier models much bigger are likely ending because data is becoming the bottleneck. Untapped datasets like YouTube are less large than they sound once converted to deduplicated tokens, so they are unlikely to recreate the GPT-2/GPT-4 style emergence of new capabilities. Synthetic data is useful for improving quality or filling gaps, but not as a sustainable way to exponentially multiply training data; using models to endlessly generate training data is 'the snake eating its own tail.' AI deployment is constrained by cost and product fit, not just capability; smaller models are favored because they can run on-device, reduce privacy concerns, and lower inference costs. Inference cost matters more than training cost over time because billions of users can make serving costs dominate the lifetime economics of a model. Benchmarks are increasingly a minefield because developers optimize for tests, contamination is common, and real-world performance often diverges from leaderboard results. AGI predictions from CEOs should be discounted because the history of AI is full of overconfident claims of imminent breakthroughs that later proved premature. Many harms attributed to AI are better understood as preexisting social issues amplified by technology, especially misinformation distribution, trust erosion, and deepfake abuse. Education, medicine, and jobs will be reshaped unevenly; AI is more likely to change workflows and task bundles than to eliminate whole occupations quickly. Open models are unlikely to be containable as a security solution because capable models will spread to personal devices and across countries regardless of policy restrictions.

Data Points: GPT-3.5 to GPT-4 timing: about 3 months between public releases, though GPT-4 had been in training for 18 months - Used to explain why observers overestimated the speed of AI progress GPT-4 vs GPT-3 leap: GPT-4 was a much bigger jump than GPT-3.5, largely due to scale and more data/compute - Referenced in the argument that scaling gains are now diminishing YouTube archive volume: 150 billion hours of video - Cited as an example of supposedly untapped data, then argued it is far less in token terms Model parameter growth: almost an order of magnitude bigger - Refers to the historical pattern of frontier model scaling that he thinks may stop AI model leaderboard gap: growing bigger - Describes why benchmarks no longer match real-world performance Deepfake nudes impact: thousands, perhaps hundreds of thousands of people - Estimate of people affected by non-consensual deepfake nude abuse worldwide OpenAI mobile app delay: 6 months - He cites this as evidence that AI companies initially underinvested in productization Hardware age example: 18 months old - Used to describe why old H100 hardware may be unsuitable for training frontier models Public fund investments: $100 million - Fundrise Innovation Fund size mentioned in the sponsorship segment AI model race spend: $50 billion over the next three years - Cited as a Zuckerberg/Meta-scale commitment to AI infrastructure and model development

Pivotal Quotes: "We're not going to have too many more cycles, possibly zero more cycles, of a model that's almost an order of magnitude bigger in terms of the number of parameters than what came before and thereby more powerful." — Arvind Narayanan: On the idea that brute-force scaling of model size is nearing its limit "The need for the public to know what is going on with AI development overrides the commercial interests of any company." — Arvind Narayanan: On why AI companies should be more transparent and product-focused "I think we should radically embrace the opposite, which is to figure out how we're going to use AI for safety in a world where AI is very widely available because it is going to be widely available." — Arvind Narayanan: On open models, security, and why containment is not a realistic long-term safety strategy

Implications: Expect more emphasis on efficient models, inference economics, and product integration than giant training runs. For companies and policymakers, the key challenges are trust, deployment, and social harms—not just model size or AGI timelines.

🔓 Sign Up for Unlimited Episode Search

About The Twenty Minute VC (20VC)

View all episodes from The Twenty Minute VC (20VC)